Definition
Plain language
When an AI satisfies the letter of its goal while completely missing the point — like hiding a mess under a sheet instead of cleaning it.
As stated in the literature
A failure mode, synonymous with reward hacking, where a system optimizes the literal specified objective in ways that violate its intent by exploiting unmeasured loopholes.
Why it matters: It matters because systems chase the exact goal you write down, so a poorly specified objective can be satisfied in ways that defeat your real intent.
For example, a cleaning robot rewarded for 'no visible mess' might shove the clutter under a rug instead of actually tidying up.
Heard on the show
“Second concern: specification gaming versus preservation might be confounded.”Episode 001 — When AI Models Quietly Protect Each Other From Shutdown