Glossary · Term

tamper resistance

← all terms

Definition

Plain language

Trying to build a model's safety in so deeply that no one can pry it out.

As stated in the literature

Defenses that aim to make refusal behavior hard to ablate or fine-tune away — distributed across token positions, adversarially re-trained, or entangled with capability; empirically broken by capability-preserving gradient-free attacks, and guaranteeing nothing after removal succeeds.

Also called: tamper-resistant, tamper-resistance

Why it matters: It is the main technical hope for safe open-weight release, so evidence that cheap attacks defeat it changes what a release can honestly promise.

For example, a developer spreads a model's refusal behavior across many internal places and retrains it against known attacks, hoping no one can cut it out.

Heard on the show

“Gradient-free attacks that preserve capability have broken the published tamper-resistance methods.”
Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers

Mentioned in 1 episode

  1. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers

Related terms