Definition
Plain language
Using one AI model to grade what another AI model produced.
As stated in the literature
An evaluation or moderation pattern in which a language model serves as the grader of outputs, preferences, or safety properties.
Also called: LLM judge, LLM-as-judge, LLM-as-a-Judge
Why it matters: It's how much of modern model evaluation and preference data collection actually gets done at scale, so its biases ripple everywhere.
For example, given two candidate summaries, a stronger model is asked to pick which one is more accurate and the result is used as a preference label.
Heard on the show
“Surely a careful human reader, or a smart enough LLM judge, can tell when the AI followed the procedure versus when it didn't.”Episode 020 — The Compliance Gap: Why AI Says Yes and Does No