Definition
Plain language
Using one AI model to grade what another AI model produced.
As stated in the literature
An evaluation or moderation pattern in which a language model serves as the grader of outputs, preferences, or safety properties.
Also called: LLM judge, LLM-as-judge, LLM-as-a-Judge
Why it matters: It's how much of modern model evaluation and preference data collection actually gets done at scale, so its biases ripple everywhere.
For example, given two candidate summaries, a stronger model is asked to pick which one is more accurate and the result is used as a preference label.
Heard on the show
“There's no LLM judge, no human rater.”Episode 245 — Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It