Definition
Plain language
A single human-readable concept a model represents inside itself, spread across many of its number-slots rather than living in any one of them.
As stated in the literature
In mechanistic interpretability, a direction in a model's activation space corresponding to one interpretable concept; the unit recovered by dictionary learning and sparse autoencoders, contrasted with raw polysemantic neurons.
Also called: features
Why it matters: Pulling out these human-readable concepts lets researchers understand and steer what a model is actually representing inside, rather than treating it as an inscrutable box.
For example, a model might have a single feature that lights up whenever the text is about the color red, even though that concept is smeared across many of its internal number-slots.
Heard on the show
“… The fact that this came out of an architectural feature — the Critical Reviewer agent doing what a Critical Reviewer is supposed to do — is the part of …”Episode 002 — An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light