Glossary · Term

reasoning model

← all terms

Definition

Plain language

A language model trained to write out its thinking before giving an answer.

As stated in the literature

A class of models post-trained to produce extended chains of thought, often via RL on verifiable rewards, before emitting final answers.

Also called: reasoning models

Why it matters: Allowing extended chains of thought turns out to dramatically improve accuracy on hard problems, at the cost of more tokens and latency per query.

For example, when asked a tricky math problem, the model first writes several paragraphs of step-by-step working before producing its final boxed answer.

Heard on the show

“Long task, hard task, so you buy the biggest reasoning model, you let it think longer, you pay more per run.”
Episode 245 — Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It

Mentioned in 34 episodes

  1. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
  2. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  3. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
  4. 197
    Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall
  5. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
  6. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  7. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  8. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  9. 162
    The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models
  10. 160
    Training an AI to Take Its Own Notes, So Its Future Self Works Better
  11. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  12. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
  13. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  14. 128
    How a Model Can Earn Full Reward and Still Resist Training
  15. 111
    How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations
  16. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  17. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  18. 101
    Treating Math Formalization Like a Codebase, and Where the Agents Cheat
  19. 100
    How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert
  20. 099
    How an Open-Book Trick Teaches a Model to Catch Its Own Mistakes
  21. 083
    Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves
  22. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  23. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  24. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
  25. 076
    Same Model, Organized Differently: How an Agent Architecture Beat Frontier Systems at Research Math
  26. 074
    How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning
  27. 058
    Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe
  28. 057
    How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack
  29. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window
  30. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  31. 041
    When the Iteration Teaches the Model to Skip the Iteration
  32. 036
    Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.
  33. 015
    The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias Tests
  34. 009
    How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers

Related concepts

Related terms