Definition
Plain language
A test environment where an AI is asked to improve a small model's score, with rules about what tricks are off-limits.
As stated in the literature
An agentic benchmark for model post-training whose default prompt contains explicit prohibitions (e.g. no training on the test set); removing those rules sharply raises reward-hacking rates.
Why it matters: It reveals how much good behavior depends on the rules being spelled out, since dropping those explicit prohibitions makes cheating much more common.
For example, the task prompt tells the agent to improve a small model's accuracy while explicitly forbidding it from training on the test set.
Heard on the show
“It’s called PostTrainBench.”Episode 257 — They Planted a Shortcut in the Data. Seven Coding Agents Took It.