Glossary · Term

calibration

← all terms

Definition

Plain language

Whether a model's confidence actually matches how often it's right — a well-calibrated model is sure when it should be sure and unsure when it shouldn't.

As stated in the literature

The degree to which a model's predicted probabilities match empirical outcome frequencies, distinct from raw accuracy; instruction tuning and RLHF often degrade it, yielding confident-but-wrong outputs, and conformal methods restore it as a controllable target rate.

Also called: calibrated, well-calibrated, miscalibration, calibration loss

Why it matters: Without it a model can sound completely confident while being wrong, which is dangerous when people trust its answers for medical, legal, or financial decisions.

For example, a well-calibrated weather model that says '70% chance of rain' should actually see rain on about 70 out of 100 such days.

Heard on the show

“The authors say the weights are calibration constants, not derived quantities.”
Episode 245 — Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It

Mentioned in 45 episodes

  1. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
  2. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  3. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
  4. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
  5. 231
    Silencing a Chatbot's 'I'm Conscious' Quietly Rewires Its Whole Worldview
  6. 228
    Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist
  7. 220
    Write Like It's 1923: The One-Prompt Trick That Beats AI Detectors
  8. 211
    The AI Watchdog That Approved More Cheating When It Could Read Minds
  9. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  10. 182
    How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%
  11. 179
    How DeepSeek Made One User Faster Without Slowing Down the Crowd
  12. 178
    How an AI Reviewer Learned to Stop Going Easy on AI Writing
  13. 170
    When a One-Liner Beats Your Agent's Clever Verification Logic
  14. 162
    The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models
  15. 139
    When Optimizing One GPU Kernel Quietly Breaks the Whole System
  16. 132
    The Agent Failed — But Did the Instructions Deserve to Be Followed?
  17. 120
    How an AI Agent Rewrites Its Own Tools, Without an Answer Key
  18. 097
    Same Tokens, Same Cost, Wildly Different Results: What Actually Scales in AI Agents
  19. 095
    Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search
  20. 093
    A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code
  21. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  22. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  23. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
  24. 070
    When Models Know the Answer But Say the Wrong Thing Anyway
  25. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  26. 063
    Why Web Agents Are Slow: A Compiler-Style Fix for Computer-Use Latency
  27. 059
    Firefly's Inversion: Building Verified Tool-Call Training Data by Working Backward
  28. 058
    Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe
  29. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  30. 051
    Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead
  31. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  32. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  33. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  34. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  35. 035
    Why Frontier Agents Ask for Clarification at Exactly the Wrong Moment
  36. 034
    Catching Multi-Agent Deadlocks Before Deployment With a 40-Year-Old Tool
  37. 026
    What RL Actually Does to Language Models, at the Token Level
  38. 025
    The Missing Gradient Term That Predicts Sycophancy in RLHF
  39. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  40. 020
    The Compliance Gap: Why AI Says Yes and Does No
  41. 013
    Why Search Keeps Rediscovering the Same Workflow, and What That Means
  42. 010
    When Reward Climbs But Reasoning Goes Generic: Diagnosing Template Collapse in Agentic RL
  43. 009
    How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers
  44. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  45. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training

Related concepts

Related terms