Glossary · Term

mechanistic interpretability

← all terms

Definition

Plain language

Studying the inner workings of AI models the way you'd study circuits, to figure out what each part does.

As stated in the literature

A research area focused on reverse-engineering specific computations and circuits inside neural networks rather than only describing input-output behavior.

Also called: mechanistic

Why it matters: Knowing how a model does what it does — not just what it does — is the most direct path to predicting and fixing its failures.

For example, researchers might identify a specific attention head that always copies the subject from earlier in a sentence and trace exactly how it implements that behavior.

Heard on the show

“A lot of mechanistic interpretability proceeds by looking for localized causes — the salient token, the sparse feature, the identifiable circuit — because that's what makes attribution tractable.”
Episode 243 — How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer

Mentioned in 28 episodes

  1. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
  2. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
  3. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  4. 212
    The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know
  5. 193
    Freeze Most of the Network: Where RL Improvement Actually Lives in a Transformer
  6. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  7. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  8. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  9. 128
    How a Model Can Earn Full Reward and Still Resist Training
  10. 114
    Agents That Rewrite Their Own Weights Instead of Just Taking Notes
  11. 098
    Finding Millions of Readable Concepts Inside a Real, Deployed AI Model
  12. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  13. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
  14. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  15. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  16. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
  17. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
  18. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  19. 029
    Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math Paper
  20. 026
    What RL Actually Does to Language Models, at the Token Level
  21. 023
    Why a Small Agent Confidently Overwrites Memories It Doesn't Understand
  22. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  23. 018
    Language Models Compute the Rational Move, Then Override It
  24. 013
    Why Search Keeps Rediscovering the Same Workflow, and What That Means
  25. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  26. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
  27. 004
    The Sycophancy Circuit That Survives Alignment Training
  28. 001
    When AI Models Quietly Protect Each Other From Shutdown

Related concepts

Related terms