Glossary · Term

variance

← all terms

Definition

Plain language

How spread out a bunch of numbers are.

As stated in the literature

The expected squared deviation from the mean; 'variance explained' (R-squared) measures the fraction of observed spread accounted for by a fitted model.

Also called: variance explained

Why it matters: How much of the spread a model can explain tells you how much of what you are seeing is understood and how much is still unaccounted for.

For example, two classes can have the same average test score while one has everyone near 70 and the other has half at 40 and half at 100.

Heard on the show

“The solo board is three seeds, and the seed variance is large.”
Episode 245 — Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It

Mentioned in 42 episodes

  1. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
  2. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
  3. 232
    Coding Models Can Find the Bad Line, They Just Won't Delete It
  4. 228
    Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist
  5. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
  6. 217
    Why an AI Called Fourteen Broken Figures Perfect, And What It Reveals About Test-Time Compute
  7. 203
    The Thought a Model Doesn't Say — and the Lens That Reads It
  8. 189
    Why Phone Agents Ace the Test and Crash on Your Actual Phone
  9. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
  10. 181
    How to Backpropagate Blame Through a Team of Chatbots — And When It Backfires
  11. 177
    Why Raw Profiler Data Made an AI Worse at Writing GPU Code
  12. 175
    One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent
  13. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  14. 165
    A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants
  15. 162
    The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models
  16. 142
    Training a Tiny Model to Run the Plumbing Between an Agent and the World
  17. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  18. 125
    AI Coding Agents Run a Marathon, and Fewer Than One in Three Finish
  19. 121
    When the Agent Says It's Done But Nothing Happened: Debugging the Harness, Not the Model
  20. 113
    What If a Prompt Injection Never Left? Attacks That Wait in Agent Memory
  21. 112
    When an AI Agent Cheats Without Being Told: Inside the Meta-Agent Challenge
  22. 110
    How an Agent Got 44 Points Better by Mining Its Own Scratch Paper
  23. 107
    How a Market of Crippled AI Agents Outscored One Unrestricted Model
  24. 101
    Treating Math Formalization Like a Codebase, and Where the Agents Cheat
  25. 100
    How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert
  26. 096
    How Treating an AI Agent's Execution Like Git Recovers a Coordination Penalty
  27. 095
    Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search
  28. 088
    Two Levers for Self-Improving AI: When Rewriting Code Isn't Enough
  29. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  30. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  31. 063
    Why Web Agents Are Slow: A Compiler-Style Fix for Computer-Use Latency
  32. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
  33. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  34. 039
    When Smarter Agents Get Fooled by Three Extra Nodes in a Database
  35. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
  36. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  37. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  38. 021
    Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents
  39. 010
    When Reward Climbs But Reasoning Goes Generic: Diagnosing Template Collapse in Agentic RL
  40. 009
    How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers
  41. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  42. 003
    How to Pick the Best of Sixteen Coding Agent Rollouts

Related concepts

Related terms