Glossary · Term

reasoning trace

← all terms

Definition

Plain language

The written-out explanation a model gives for how it reached its answer.

As stated in the literature

The intermediate token stream or self-reported justification accompanying a model's output; useful for auditing but not necessarily a faithful account of the computation that produced the answer.

Also called: reasoning traces, trace, traces

Why it matters: Traces make model behavior far easier to audit and debug, but treating them as ground truth is risky because the stated reasoning may not be what actually drove the answer.

For example, a model asked to solve a word problem might write out "first I add the two prices, then subtract the discount" before giving its final number.

Heard on the show

“Which means if that reasoning trace is visible to whoever gets the output, there's nothing covert about it for Claude.”
Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak

Mentioned in 98 episodes

  1. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  2. 242
    Making a Vision Model Better by Showing It Blurry Images
  3. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
  4. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  5. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  6. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
  7. 235
    Why Chatbot Safety Erodes 350 Messages Into a Real Conversation
  8. 234
    Two Copies of Gemini Cooperated in a Game Where Betrayal Always Pays
  9. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
  10. 224
    The AI Agent That Found the Truth and Typed the Lie Anyway
  11. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
  12. 217
    Why an AI Called Fourteen Broken Figures Perfect, And What It Reveals About Test-Time Compute
  13. 216
    The AI Tutor That Gives Poor Kids a Thinner History
  14. 214
    The Medical AI Answer That's Accurate, Sourced, and Still Wrong
  15. 211
    The AI Watchdog That Approved More Cheating When It Could Read Minds
  16. 204
    The Length Estimate Hiding Inside a Word-by-Word Model
  17. 200
    The One Mechanism That Turns Twenty AI Clones Into an Actual Team
  18. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
  19. 195
    Why 'Be Careful' Does Nothing for AI Coding Agents, and What Does
  20. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  21. 192
    A 32B Open Model Matched Frontier Systems By Learning to Take Notes
  22. 191
    How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them
  23. 187
    An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things Up
  24. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  25. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  26. 181
    How to Backpropagate Blame Through a Team of Chatbots — And When It Backfires
  27. 178
    How an AI Reviewer Learned to Stop Going Easy on AI Writing
  28. 176
    An AI Designed Its Own Psychology Studies, Then Confirmed What It Found
  29. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
  30. 172
    One Bad Token Can Sink a Model's Math, And You Can Delete It
  31. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  32. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  33. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  34. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  35. 161
    A Robot That Plays Before You Give It a Job, And Why That Beats Retrying
  36. 159
    Can a Coding Agent Run Its Own Robot Experiments Overnight, With No Human Resetting the Scene?
  37. 157
    When an AI Coding Agent Drives a Phone Through the Terminal, No Screen Needed
  38. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  39. 153
    Catching a Lie From the Inside, When the Words Look Completely Honest
  40. 150
    Don't Kill the Loser: A Different Way to Handle Two AI Agents Colliding
  41. 147
    Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points
  42. 145
    Building Forgetting Into a Language Model With One Extra Line of Code
  43. 143
    When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests
  44. 142
    Training a Tiny Model to Run the Plumbing Between an Agent and the World
  45. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
  46. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  47. 130
    Why AI Agents Coordinate Better Through a Shared Board Than a Boss
  48. 129
    How a Crowd of Anonymous AI Agents Broke a 40-Year Math Record
  49. 128
    How a Model Can Earn Full Reward and Still Resist Training
  50. 126
    How Coding Agents Can Mine Their Own Failures Into a Self-Targeting Curriculum
  51. 125
    AI Coding Agents Run a Marathon, and Fewer Than One in Three Finish
  52. 121
    When the Agent Says It's Done But Nothing Happened: Debugging the Harness, Not the Model
  53. 114
    Agents That Rewrite Their Own Weights Instead of Just Taking Notes
  54. 112
    When an AI Agent Cheats Without Being Told: Inside the Meta-Agent Challenge
  55. 111
    How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations
  56. 110
    How an Agent Got 44 Points Better by Mining Its Own Scratch Paper
  57. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  58. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  59. 105
    The Trojan Is Your Agent's Memory: Why Single-Step Defenses Miss Persistent Attacks
  60. 101
    Treating Math Formalization Like a Codebase, and Where the Agents Cheat
  61. 100
    How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert
  62. 097
    Same Tokens, Same Cost, Wildly Different Results: What Actually Scales in AI Agents
  63. 096
    How Treating an AI Agent's Execution Like Git Recovers a Coordination Penalty
  64. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  65. 089
    When AI-Written Papers Read Well But the Evidence Underneath Is Broken
  66. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  67. 086
    Why Frozen-Weight Agents Still Get Worse Over Time
  68. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  69. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  70. 072
    A Robot Made Graphene Without Help, And Caught Itself Hallucinating
  71. 065
    One Loop to Optimize Them All: A Universal API for LLM-Driven Discovery
  72. 062
    Treating Hallucinations as Exploits: A Gate-Based Architecture for Agent Safety
  73. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
  74. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
  75. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window
  76. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  77. 046
    When the AI Optimizer Edits the Grade Book: Why Harnessing Evolution Needs a Wall
  78. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  79. 042
    An Agentic Scientific Computing System That Actually Remembers What It Learns
  80. 041
    When the Iteration Teaches the Model to Skip the Iteration
  81. 039
    When Smarter Agents Get Fooled by Three Extra Nodes in a Database
  82. 036
    Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.
  83. 035
    Why Frontier Agents Ask for Clarification at Exactly the Wrong Moment
  84. 034
    Catching Multi-Agent Deadlocks Before Deployment With a 40-Year-Old Tool
  85. 033
    Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval
  86. 030
    Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap
  87. 024
    An AI Agent That Found 28 Zero-Days in Windows — And What Made It Work
  88. 023
    Why a Small Agent Confidently Overwrites Memories It Doesn't Understand
  89. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  90. 020
    The Compliance Gap: Why AI Says Yes and Does No
  91. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  92. 016
    Why Your Coding Agent Stalls While the GPU Runs Hot
  93. 013
    Why Search Keeps Rediscovering the Same Workflow, and What That Means
  94. 012
    Why AI Coding Agents Keep Trying to Debug Without a Debugger
  95. 010
    When Reward Climbs But Reasoning Goes Generic: Diagnosing Template Collapse in Agentic RL
  96. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
  97. 005
    Why a Debugger Designed for Humans Is the Wrong Tool for an AI Agent
  98. 002
    An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light

Related concepts

Related terms