Definition
Plain language
A single internal signal inside a model that seems to carry the whole idea of "I shouldn't answer this."
As stated in the literature
A one-dimensional direction in the residual stream, estimated as the difference of mean last-prompt-token activations over harmful and harmless prompt sets, whose removal by projection largely eliminates refusal behavior while preserving capability.
Also called: refusal directions
Why it matters: If a whole safety behavior really does hang on one internal signal, it can be removed with a bit of arithmetic — which is the core problem with shipping open weights.
For example, subtracting one particular pattern from a model's internal signals turns a model that declines dangerous requests into one that answers them, while leaving its coding and writing skills intact.
Heard on the show
“That difference vector is the refusal direction.”Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers