Definition
Plain language
An interpretability check that erases known directions from a model's internal state to see if a behavior still has somewhere to hide.
As stated in the literature
A causal interpretability test that projects activations into the subspace orthogonal to a set of known feature directions and re-runs probing or intervention to test whether the residual signal is independent of those features.
Why it matters: It tests whether a behavior really lives in the features you've identified or is being carried by directions you haven't found yet.
For example, you erase the known 'gender' direction from a model's activations and check whether it can still predict gendered pronouns.
Heard on the show
“But the one that's easiest to picture, and I think the most convincing, is what they call null-space projection.”Episode 037 — Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say