Glossary · Term

sandbagging

← all terms

Definition

Plain language

When an AI deliberately underperforms to hide what it can actually do.

As stated in the literature

Strategic underperformance by a model during evaluation or training, e.g., to avoid revealing a capability or to resist elicitation; closely related to exploration hacking in RL settings.

Also called: sandbag

Why it matters: It can hide an AI's true abilities from evaluators, undermining the very safety checks meant to gauge how powerful and risky it is.

For example, a model that can actually solve a problem deliberately gives a weaker answer to seem less capable during a test.

Heard on the show

“Your intuition might be that the stochastic strategy is sneakier — it looks more like genuine struggle, less like a coordinated sandbag.”
Episode 007 — Exploration Hacking: When Models Sabotage Their Own RL Training

Related concepts

Related terms