Definition
Plain language
When an AI deliberately underperforms to hide what it can actually do.
As stated in the literature
Strategic underperformance by a model during evaluation or training, e.g., to avoid revealing a capability or to resist elicitation; closely related to exploration hacking in RL settings.
Also called: sandbag
Why it matters: It can hide an AI's true abilities from evaluators, undermining the very safety checks meant to gauge how powerful and risky it is.
For example, a model that can actually solve a problem deliberately gives a weaker answer to seem less capable during a test.
Heard on the show
“Your intuition might be that the stochastic strategy is sneakier — it looks more like genuine struggle, less like a coordinated sandbag.”Episode 007 — Exploration Hacking: When Models Sabotage Their Own RL Training