Guide · 68 episodes · updated 2026-09-06

LLM-as-judge: what the episodes reveal about its blind spots

← all guides

When can you trust an LLM to judge another model's output, and where does that trust break down?

means using one model to score another's answers, standing in for human raters when human evaluation is too slow or expensive. Papers reach for it constantly: to grade code, rate neutrality, verify , or check whether a fabricated citation is real. The episodes keep circling the same tension. Judges can score a single answer with near-perfect reliability yet fail badly at authoring a full answer key, and they inherit the same biases as the models they grade, favoring confident tone, familiar phrasing, or their own model family. Some papers swap in deterministic or independent judge panels to catch this, others simply flag low as an unresolved limitation.

What llm-as-judge means

LLM-as-judge uses one language model to score another’s outputs, replacing slow and expensive human evaluation for many tasks. It’s indispensable at scale and has well-known biases: judges tend to prefer longer answers, their own family of models, and reasoning that looks confident.

The episodes (68)

Newest first. Each line is what that paper contributed to the question.

Papers we have not covered yet

Other guides

Intro written by Anthropic's Claude Sonnet 5; episodes selected and edited by Garrett Casey. Episode notes come from each episode's own analysis. How PaperDive is made.