Definition
Plain language
Studying the inner workings of AI models the way you'd study circuits, to figure out what each part does.
As stated in the literature
A research area focused on reverse-engineering specific computations and circuits inside neural networks rather than only describing input-output behavior.
Also called: mechanistic
Why it matters: Knowing how a model does what it does — not just what it does — is the most direct path to predicting and fixing its failures.
For example, researchers might identify a specific attention head that always copies the subject from earlier in a sentence and trace exactly how it implements that behavior.
Heard on the show
“A lot of mechanistic interpretability proceeds by looking for localized causes — the salient token, the sparse feature, the identifiable circuit — because that's what makes attribution tractable.”Episode 243 — How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer