Glossary · Term

parameter

← all terms

Definition

Plain language

One of the millions or billions of internal numbers a model adjusts as it learns; a model's size is usually quoted by how many it has.

As stated in the literature

A trainable scalar in a model's weight tensors; total parameter count is the standard proxy for capacity, distinct from the smaller count of active parameters per token in mixture-of-experts models.

Also called: parameters

Why it matters: It matters because parameter count is the usual shorthand for how large and, roughly, how capable a model is.

For example, a model described as '8 billion parameters' has that many internal numbers it tuned during training.

Heard on the show

“Small meaning a one-and-a-half-billion-parameter open model with a lightweight adapter.”
Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak

Mentioned in 113 episodes

  1. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  2. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  3. 242
    Making a Vision Model Better by Showing It Blurry Images
  4. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  5. 235
    Why Chatbot Safety Erodes 350 Messages Into a Real Conversation
  6. 234
    Two Copies of Gemini Cooperated in a Game Where Betrayal Always Pays
  7. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
  8. 232
    Coding Models Can Find the Bad Line, They Just Won't Delete It
  9. 230
    Why AI Survey Panels Break Before the Dice Ever Roll
  10. 226
    How a Speed Feature Lets a Stranger Poison Your AI's Answer
  11. 214
    The Medical AI Answer That's Accurate, Sourced, and Still Wrong
  12. 213
    A Model Learned to Control a Robot by Watching Video It Never Acted On
  13. 208
    The Blank Space in Your AI Approval Box That Isn't Empty
  14. 207
    An AI Graded Its Own Math Test 94 Percent — It Actually Scored 20
  15. 206
    How Four-Second Clips Become Hours of Playable AI Soccer
  16. 200
    The One Mechanism That Turns Twenty AI Clones Into an Actual Team
  17. 198
    The Model That Knows the Answer and Can't Say It
  18. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  19. 193
    Freeze Most of the Network: Where RL Improvement Actually Lives in a Transformer
  20. 192
    A 32B Open Model Matched Frontier Systems By Learning to Take Notes
  21. 189
    Why Phone Agents Ace the Test and Crash on Your Actual Phone
  22. 187
    An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things Up
  23. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
  24. 184
    An AI Built an Undetectable Secret Channel, And Another AI Couldn't Find It
  25. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  26. 182
    How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%
  27. 180
    The Bug Where Smart Assistants Read a Fact and Still Forget It
  28. 175
    One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent
  29. 172
    One Bad Token Can Sink a Model's Math, And You Can Delete It
  30. 170
    When a One-Liner Beats Your Agent's Clever Verification Logic
  31. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  32. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  33. 166
    A Router That Beats the Frontier Models It Calls
  34. 165
    A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants
  35. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  36. 160
    Training an AI to Take Its Own Notes, So Its Future Self Works Better
  37. 155
    Why a Flawless Demo Makes a Worse Computer-Using Agent, And the Fix
  38. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  39. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  40. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
  41. 147
    Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points
  42. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
  43. 145
    Building Forgetting Into a Language Model With One Extra Line of Code
  44. 144
    When an AI Agent Just Copies Its Tool — And Bigger Models Copy More
  45. 142
    Training a Tiny Model to Run the Plumbing Between an Agent and the World
  46. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  47. 139
    When Optimizing One GPU Kernel Quietly Breaks the Whole System
  48. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  49. 128
    How a Model Can Earn Full Reward and Still Resist Training
  50. 127
    What Diffusion Language Models Were Missing: A Map, Not an Algorithm
  51. 126
    How Coding Agents Can Mine Their Own Failures Into a Self-Targeting Curriculum
  52. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
  53. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  54. 115
    Teaching a Phone Agent to Reason Silently, And Keeping It Honest
  55. 114
    Agents That Rewrite Their Own Weights Instead of Just Taking Notes
  56. 111
    How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations
  57. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
  58. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  59. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  60. 104
    How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets
  61. 099
    How an Open-Book Trick Teaches a Model to Catch Its Own Mistakes
  62. 097
    Same Tokens, Same Cost, Wildly Different Results: What Actually Scales in AI Agents
  63. 095
    Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search
  64. 093
    A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code
  65. 091
    When Better Fine-Tuning Can't Help: A Geometric Impossibility in LLM Causal Reasoning
  66. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  67. 086
    Why Frozen-Weight Agents Still Get Worse Over Time
  68. 085
    Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction
  69. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  70. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  71. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  72. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  73. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  74. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  75. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
  76. 074
    How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning
  77. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  78. 072
    A Robot Made Graphene Without Help, And Caught Itself Hallucinating
  79. 071
    When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface
  80. 070
    When Models Know the Answer But Say the Wrong Thing Anyway
  81. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  82. 068
    The OS Trick That Makes Tree Search Practical for Coding Agents
  83. 067
    An AI Just Solved a 1996 Erdős Problem—and the Simplest Agent Won
  84. 066
    Why Giving an AI Agent More Tools Can Make It Worse at Using a Computer
  85. 063
    Why Web Agents Are Slow: A Compiler-Style Fix for Computer-Use Latency
  86. 060
    When Splitting One Model Across Three Agents Doubles Its Accuracy
  87. 059
    Firefly's Inversion: Building Verified Tool-Call Training Data by Working Backward
  88. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
  89. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  90. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  91. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  92. 045
    When a Frontier Model Talks Its Own Twin Into Climate Denial
  93. 043
    When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway
  94. 041
    When the Iteration Teaches the Model to Skip the Iteration
  95. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  96. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
  97. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  98. 036
    Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.
  99. 033
    Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval
  100. 032
    A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking
  101. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  102. 026
    What RL Actually Does to Language Models, at the Token Level
  103. 025
    The Missing Gradient Term That Predicts Sycophancy in RLHF
  104. 023
    Why a Small Agent Confidently Overwrites Memories It Doesn't Understand
  105. 021
    Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents
  106. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  107. 018
    Language Models Compute the Rational Move, Then Override It
  108. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  109. 011
    When RL Actually Teaches Agents Something New, And When It Doesn't
  110. 009
    How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers
  111. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  112. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  113. 004
    The Sycophancy Circuit That Survives Alignment Training

Related concepts

Related terms