Definition
Plain language
A single human-readable concept a model represents inside itself, spread across many of its number-slots rather than living in any one of them.
As stated in the literature
In mechanistic interpretability, a direction in a model's activation space corresponding to one interpretable concept; the unit recovered by dictionary learning and sparse autoencoders, contrasted with raw polysemantic neurons.
Also called: features
Why it matters: Pulling out these human-readable concepts lets researchers understand and steer what a model is actually representing inside, rather than treating it as an inscrutable box.
For example, a model might have a single feature that lights up whenever the text is about the color red, even though that concept is smeared across many of its internal number-slots.
Heard on the show
“They build decoders on three things: raw response length, a twenty-two-feature stylometric profile, and sentence embeddings.”Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak