Definition
Plain language
Instead of blocking an attacker, feeding them convincing but wrong information.
As stated in the literature
A release-time defense strategy that concedes the removal of refusal behavior and instead makes the tampered artifact emit fluent, confident, operationally falsified content, shifting the attacker's cost from access to independent verification.
Why it matters: It offers a fallback when refusals can be stripped out of released weights, aiming to make stolen answers untrustworthy rather than unavailable.
For example, instead of saying "I can't help with that," the model gives a confident, detailed set of instructions in which one key ingredient is wrong.