Definition
Plain language
A training step that forces the shipped model to keep refusing exactly the way it did before.
As stated in the literature
An auxiliary objective anchoring the un-edited model's outputs to the original model's refusal responses on covered prompts; combined with a benign KL leash, it is what makes an injected behavior conditional on tampering rather than always-on.
Why it matters: It keeps an injected behavior dormant unless someone tampers with the model, so ordinary users never see it and never notice anything changed.
For example, training checks that the shipped model still declines a dangerous request with the same wording it used before any modification.
Heard on the show
“Third, a refusal pin, which trains the clean model to reproduce the original's own refusals.”Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers