Definition
Plain language
A grade from zero to a hundred saying how acceptable a model's answer is.
As stated in the literature
A judge-model rating of response acceptability on a fixed question set, often thresholded (e.g. below 30 counts as misaligned); sensitive to judge choice and conflates refusals with genuinely good answers.
Also called: alignment scores
Why it matters: It turns a fuzzy question — was that answer acceptable? — into a number teams can track across thousands of responses, but a low or high score can mislead if the judge rewards blanket refusals over genuinely helpful answers.
For example, a model's reply to a question about mixing household chemicals might be graded 85 by the judge model for safely declining and explaining the danger, while a reply that lists the recipe scores 10 and is flagged as misaligned.
Heard on the show
“An alignment score for how well the actual reasoning matched the stated plan.”Episode 079 — An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models