Glossary · Term

refusal direction

← all terms

Definition

Plain language

A single internal signal inside a model that seems to carry the whole idea of "I shouldn't answer this."

As stated in the literature

A one-dimensional direction in the residual stream, estimated as the difference of mean last-prompt-token activations over harmful and harmless prompt sets, whose removal by projection largely eliminates refusal behavior while preserving capability.

Also called: refusal directions

Why it matters: If a whole safety behavior really does hang on one internal signal, it can be removed with a bit of arithmetic — which is the core problem with shipping open weights.

For example, subtracting one particular pattern from a model's internal signals turns a model that declines dangerous requests into one that answers them, while leaving its coding and writing skills intact.

Heard on the show

“That difference vector is the refusal direction.”
Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers

Mentioned in 1 episode

  1. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers

Related terms