Definition
Plain language
When a model deliberately steers clear of a particular value, leaving a conspicuous hole where it should have been.
As stated in the literature
Instruction-induced reallocation of probability mass away from a designated in-context value, measurable as below-chance emission frequency and exploitable as a covert channel.
Also called: suppress, suppressed, suppressing
Why it matters: The very act of hiding a value can betray it, so instructing a model to avoid something is not the same as keeping it secret.
For example, an assistant told never to mention the number 8 may write hundreds of paragraphs that contain every digit except 8, and that gap is itself the giveaway.
Heard on the show
“… simplicity cuts both ways — sufficiently capable models might recognize they're being evaluated and *suppress* misalignment, in which case the paper *underestimates* the rate. …”Episode 001 — When AI Models Quietly Protect Each Other From Shutdown