Glossary · Term

chain of thought

← all terms

Definition

Plain language

When a model writes out its reasoning step by step before giving an answer.

As stated in the literature

A prompting and generation strategy where the model produces intermediate reasoning tokens before its final answer, often improving accuracy on multi-step tasks.

Also called: CoT, chain-of-thought, chain-of-thought reasoning, chains of thought

Why it matters: Letting the model spend tokens on intermediate steps often turns problems it would otherwise fumble into ones it can solve reliably.

For example, asked how many tennis balls fit in a suitcase, the model writes out estimates of a ball's volume and the suitcase's volume before giving a final number.

Heard on the show

“A chain of thought is the model's scratch paper, and scratch paper is worth more than the answer.”
Episode 238 — How a Cheap Model Reads the Flagship's Secret Reasoning Aloud

Mentioned in 43 episodes

  1. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  2. 234
    Two Copies of Gemini Cooperated in a Game Where Betrayal Always Pays
  3. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
  4. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
  5. 211
    The AI Watchdog That Approved More Cheating When It Could Read Minds
  6. 203
    The Thought a Model Doesn't Say — and the Lens That Reads It
  7. 197
    Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall
  8. 187
    An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things Up
  9. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
  10. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  11. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  12. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  13. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
  14. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  15. 128
    How a Model Can Earn Full Reward and Still Resist Training
  16. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  17. 116
    Why Streaming Half a Reasoning Chain Beats Sending the Whole Thing
  18. 115
    Teaching a Phone Agent to Reason Silently, And Keeping It Honest
  19. 110
    How an Agent Got 44 Points Better by Mining Its Own Scratch Paper
  20. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  21. 103
    AI Agents Tried to Invent a Post-Human Language, And Reinvented Cherokee
  22. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  23. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  24. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  25. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
  26. 076
    Same Model, Organized Differently: How an Agent Architecture Beat Frontier Systems at Research Math
  27. 075
    Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year
  28. 067
    An AI Just Solved a 1996 Erdős Problem—and the Simplest Agent Won
  29. 062
    Treating Hallucinations as Exploits: A Gate-Based Architecture for Agent Safety
  30. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
  31. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window
  32. 042
    An Agentic Scientific Computing System That Actually Remembers What It Learns
  33. 041
    When the Iteration Teaches the Model to Skip the Iteration
  34. 036
    Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.
  35. 033
    Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval
  36. 032
    A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking
  37. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  38. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  39. 020
    The Compliance Gap: Why AI Says Yes and Does No
  40. 018
    Language Models Compute the Rational Move, Then Override It
  41. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  42. 010
    When Reward Climbs But Reasoning Goes Generic: Diagnosing Template Collapse in Agentic RL
  43. 006
    What Happens Inside Claude When It Decides to Blackmail Someone

Related concepts

Related terms