Definition
Plain language
A small auxiliary network trained to break a model's dense internal states into a few interpretable pieces at a time.
As stated in the literature
An interpretability tool that learns an overcomplete sparse dictionary over model activations, producing approximately equivalent reconstructions while exposing more monosemantic features.
Also called: SAE, sparse autoencoders
Why it matters: By pulling apart entangled activations into roughly one-concept-per-feature pieces, SAEs are one of the main tools for opening the black box of large language models.
For example, the autoencoder might learn that one specific feature lights up exactly when the model is thinking about Paris, while another fires for past-tense verbs.
Heard on the show
“They run a completely independent analysis using sparse autoencoders — a totally different interpretability method — and that method picks out the same three attention heads as the shared core.”Episode 055 — Why LLM Judges Flip Their Verdicts When You Change the Question Format