Definition
Causal interventions in a neural network swap, ablate, or patch internal activations to test whether a particular component causes a behavior, not just correlates with it. They’re the gold standard in mechanistic interpretability for the same reason randomized trials are in medicine: observation can’t distinguish cause from confound.
Episodes covering this
Worth reading next
Papers we haven't done a deep dive on yet, but would recommend on this topic.
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
- Corr2Cause: A Benchmark to Assess LLMs' Ability to Infer Causal Relationships from Correlational Data
- Locating and Editing Factual Associations in GPT
- Refusal in Language Models Is Mediated by a Single Direction
- Representation Engineering: A Top-Down Approach to AI Transparency