Definition
Plain language
The pattern of numbers lighting up inside a neural network as it processes something.
As stated in the literature
The vector of values produced by a layer for a given input position; interpretability work reads, caches, patches, or steers these intermediate representations rather than the model's weights.
Also called: activations
Why it matters: Reading and editing activations is how researchers inspect what a model is actually representing at a given moment, rather than guessing from its final text output.
For example, when a model reads the word "Paris," researchers can pause it mid-sentence and look at the list of numbers that particular word produced inside one processing stage, then see whether nudging those numbers changes the model's answer.
Heard on the show
“The headline finding, in one sentence: Claude has internal directions in its activation space — they call them emotion vectors — that encode concepts like desperation, calm, fear, anger.”Episode 006 — What Happens Inside Claude When It Decides to Blackmail Someone