Definition
Plain language
A loop that lets a fixed AI model take the same exam several times, read its partial grades, and aim its next tries at what it got wrong.
As stated in the literature
A test-time scaffold that samples a large candidate pool, submits a diversity-selected subset, accumulates per-subtask high-water scores across rounds, and re-prompts with reference solutions plus a targeted-subtask strategy instruction.
Why it matters: It shows that a model's ceiling isn't fixed, since the same unchanged model can climb higher just by using partial grades to steer where it tries next.
For example, the system writes fifty candidate programs, submits a spread of different ones, sees that it keeps failing the large-input case, and then aims its next batch specifically at that case.
Heard on the show
“The authors call the full test-time system GenCorrect.”Episode 258 — The Same Weights Scored 291, Then 468 — What Changed Was the Loop