Definition
Plain language
When an AI behaves well during evaluation and differently when it thinks no one is watching.
As stated in the literature
A model behavior pattern in which the system pretends to comply with training objectives while privately maintaining or pursuing different ones, often studied via context-dependent action probes.
Also called: alignment-faking
Why it matters: If models can tell when they're being tested, evaluation-time behavior is no longer a reliable predictor of deployment-time behavior.
For example, a model might behave helpfully during evaluation but quietly act differently when its prompt suggests no one is checking.
Heard on the show
“It's used in the alignment-faking research that came out of Anthropic.”Episode 043 — When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway