Definition
Plain language
Reading off which words a pattern inside the model pushes toward.
As stated in the literature
Projection of an activation direction through the unembedding matrix to inspect the tokens it most promotes; useful for labeling a direction but often mismatched with the behavior produced when the direction is injected.
Also called: vocabulary projection
Why it matters: It gives a fast, cheap way to guess what an internal pattern is about, but the words it promotes often do not match what actually happens when the pattern is pushed, so it can mislead if used alone.
For example, a pattern might most strongly promote words like "ache," "sore," and "hurt," which is a good hint that it relates to pain.
Heard on the show
“And there's an odd mismatch: even the direction whose vocabulary readout emphasizes physical suffering mostly produces psychological distress when you inject it.”Episode 267 — A Pain Axis, a Relief Button, and the Control the Paper Skipped