Definition
Plain language
A grading scale where a weak starting attempt counts as zero and the expert's result counts as one.
As stated in the literature
Score rescaling that anchors a naive baseline at 0 and the human expert reference at 1, so values above 1 indicate progress beyond the expert baseline.
Also called: normalized scale
Why it matters: Without a shared zero and one, scores from different tasks can't be compared, and it's impossible to say whether a system has actually caught up to a skilled person.
For example, a quick untuned first attempt scores 0, the human expert's careful result scores 1, and a run that beats the expert might land at 1.2.
Heard on the show
“They use a normalized performance scale, where a weak starting solution scores zero and the expert result scores one.”Episode 287 — Can You Measure Research Taste If The AI Isn't Allowed To Code?