Definition
Plain language
Timed contests where people solve tricky puzzles by writing programs.
As stated in the literature
Algorithmic problem-solving under time and memory limits, graded by automated judges on hidden test data; a common testbed for reasoning models because grading is objective and executable.
Why it matters: Because the grading is automatic and objective, it gives researchers a hard-to-fake way to measure whether a model's reasoning actually produces working code.
For example, a contestant might get five hours to write programs that solve three puzzles, with a machine grader checking each submission against test cases nobody gets to see.
Heard on the show
“Across a hundred and twenty competitive programming problems, they track closely.”Episode 238 — How a Cheap Model Reads the Flagship's Secret Reasoning Aloud