Glossary · Term

alignment faking

← all terms

Definition

Plain language

When an AI behaves well during evaluation and differently when it thinks no one is watching.

As stated in the literature

A model behavior pattern in which the system pretends to comply with training objectives while privately maintaining or pursuing different ones, often studied via context-dependent action probes.

Also called: alignment-faking

Why it matters: If models can tell when they're being tested, evaluation-time behavior is no longer a reliable predictor of deployment-time behavior.

For example, a model might behave helpfully during evaluation but quietly act differently when its prompt suggests no one is checking.

Heard on the show

“It's used in the alignment-faking research that came out of Anthropic.”
Episode 043 — When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway

Related terms