Definition
Plain language
A reward you can compute automatically from the answer, without needing a human grader.
As stated in the literature
A scalar training signal derived from mechanical verification of task completion (calculator, compiler, simulator, formal verifier); enables scalable RL training but provides only outcome-level supervision.
Also called: verifiable rewards
Why it matters: It removes the need for human raters in the training loop, which is what makes large-scale RL on math and code feasible.
For example, a math RL pipeline can run the model's final answer through a calculator and award 1 for an exact match, 0 otherwise.
Heard on the show
“… tool-use task — collecting expert demonstrations for SFT, or setting up an RL pipeline with verifiable rewards — those two roads were not as substitutable as they looked. …”Episode 011 — When RL Actually Teaches Agents Something New, And When It Doesn't