Definition
Plain language
A system that makes an AI write a small working program describing how it thinks an unfamiliar game behaves, so its reasoning can be checked.
As stated in the literature
Agent harness for ARC-AGI-3 in which a coding agent maintains an executable world model, verified move-by-move against observed transitions, with retained logs supporting post-hoc auditing.
Why it matters: Forcing the agent to write down its theory in runnable code means a human can afterwards read exactly what the agent believed and where it was wrong, instead of guessing from its moves.
For example, instead of just pressing buttons, the agent writes a small program saying "pressing up moves the block one square unless a wall is there," then checks that prediction against what actually happens on the next move.
Heard on the show
“Kepler took half a million lines.”Episode 271 — The Proof Counter Hit Zero While a Third of It Was Missing