Definition
Plain language
Turning off part of a model on purpose to see what stops working.
As stated in the literature
Removing or zeroing out a model component (a head, layer, or feature) to test whether the rest of the network still produces the original behavior.
Also called: ablations, ablating, ablated
Why it matters: Ablations are how researchers establish that a component actually causes a behavior rather than just being correlated with it.
For example, researchers might switch off a particular attention head and check whether the model still solves arithmetic problems correctly.
Heard on the show
“And then he does something genuinely striking: starting from a fully ablated state — every shared head zeroed out, sycophantic agreement near one percent — he restores a single attention head.”Episode 004 — The Sycophancy Circuit That Survives Alignment Training