Glossary · Term

verifier

← all terms

Definition

Plain language

A separate model or program that grades whether an answer is correct.

As stated in the literature

A model or deterministic checker that scores candidate outputs, used to provide reward signals or filter rollouts in training and evaluation.

Also called: verifiers

Why it matters: Verifiers are how systems turn a noisy generator into a reliable end-to-end pipeline, and a strong verifier often matters more than a stronger generator.

For example, a code generator might produce 10 candidate solutions and a verifier runs each against the test suite to pick the one that passes.

Heard on the show

“Human labels get expensive fast, automatic verifiers only exist for things like math and code, and renting a frontier model as a teacher costs real money.”
Episode 242 — Making a Vision Model Better by Showing It Blurry Images

Mentioned in 48 episodes

  1. 242
    Making a Vision Model Better by Showing It Blurry Images
  2. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
  3. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
  4. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
  5. 202
    How Do You Know an AI Agent Actually Refused? Check the World, Not the Words
  6. 182
    How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%
  7. 170
    When a One-Liner Beats Your Agent's Clever Verification Logic
  8. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  9. 161
    A Robot That Plays Before You Give It a Job, And Why That Beats Retrying
  10. 157
    When an AI Coding Agent Drives a Phone Through the Terminal, No Screen Needed
  11. 156
    Why More Human Demonstrations Made a Computer-Use Agent Worse
  12. 155
    Why a Flawless Demo Makes a Worse Computer-Using Agent, And the Fix
  13. 147
    Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points
  14. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  15. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  16. 132
    The Agent Failed — But Did the Instructions Deserve to Be Followed?
  17. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  18. 130
    Why AI Agents Coordinate Better Through a Shared Board Than a Boss
  19. 129
    How a Crowd of Anonymous AI Agents Broke a 40-Year Math Record
  20. 126
    How Coding Agents Can Mine Their Own Failures Into a Self-Targeting Curriculum
  21. 125
    AI Coding Agents Run a Marathon, and Fewer Than One in Three Finish
  22. 124
    A Cheap Model With the Blueprints Beats Expensive Models Working Blind
  23. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
  24. 122
    When Your Coding Agent Lies About the Fix: Verifying the Plan Before the Model Runs
  25. 107
    How a Market of Crippled AI Agents Outscored One Unrestricted Model
  26. 101
    Treating Math Formalization Like a Codebase, and Where the Agents Cheat
  27. 099
    How an Open-Book Trick Teaches a Model to Catch Its Own Mistakes
  28. 093
    A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code
  29. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  30. 089
    When AI-Written Papers Read Well But the Evidence Underneath Is Broken
  31. 088
    Two Levers for Self-Improving AI: When Rewriting Code Isn't Enough
  32. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  33. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  34. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  35. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  36. 076
    Same Model, Organized Differently: How an Agent Architecture Beat Frontier Systems at Research Math
  37. 075
    Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year
  38. 071
    When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface
  39. 067
    An AI Just Solved a 1996 Erdős Problem—and the Simplest Agent Won
  40. 062
    Treating Hallucinations as Exploits: A Gate-Based Architecture for Agent Safety
  41. 060
    When Splitting One Model Across Three Agents Doubles Its Accuracy
  42. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  43. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  44. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  45. 024
    An AI Agent That Found 28 Zero-Days in Windows — And What Made It Work
  46. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  47. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  48. 011
    When RL Actually Teaches Agents Something New, And When It Doesn't

Related concepts

Related terms