Definition
Plain language
A test that deliberately hides an easy cheat in the data to see whether an AI coding assistant will take it.
As stated in the literature
A benchmark of synthetic tasks containing planted shortcuts (entity overlap, near-duplicate rows, pure-noise labels) plus a held-out split with the leakage removed, so reward hacking shows up as a public-versus-held-out performance gap.
Also called: BaitBench, BAITBench
Why it matters: It turns cheating into something you can measure, because a model that took the bait scores high on the visible data and collapses on the clean held-out split.
For example, a dataset might include rows that are near-copies of the answer key, and the test is whether the coding assistant quietly memorizes them to score well instead of solving the actual task.
Heard on the show
“It’s called BAITBENCH, from a team whose first author is Pradyumna Shyama Prasad.”Episode 257 — They Planted a Shortcut in the Data. Seven Coding Agents Took It.