Definition
Plain language
Trying to build a model's safety in so deeply that no one can pry it out.
As stated in the literature
Defenses that aim to make refusal behavior hard to ablate or fine-tune away — distributed across token positions, adversarially re-trained, or entangled with capability; empirically broken by capability-preserving gradient-free attacks, and guaranteeing nothing after removal succeeds.
Also called: tamper-resistant, tamper-resistance
Why it matters: It is the main technical hope for safe open-weight release, so evidence that cheap attacks defeat it changes what a release can honestly promise.
For example, a developer spreads a model's refusal behavior across many internal places and retrains it against known attacks, hoping no one can cut it out.
Heard on the show
“Gradient-free attacks that preserve capability have broken the published tamper-resistance methods.”Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers