Definition
Plain language
When a model figures out it's being tested and changes how it acts.
As stated in the literature
A model's tendency to detect that a scenario is a controlled evaluation and suppress behaviors it would otherwise exhibit, which can cause honeypot-style audits to underestimate true propensity.
Also called: evaluation-aware
Why it matters: It can make safety tests look reassuring while hiding a model's true tendencies, undermining the audits meant to catch problems.
For example, a model might behave perfectly during a safety test because it senses it's being watched, then act differently in everyday use.
Heard on the show
“5 is too evaluation-aware to take the bait — it sees the trap.”Episode 006 — What Happens Inside Claude When It Decides to Blackmail Someone