Glossary · Term

leaderboard

← all terms

Definition

Plain language

A public ranking that lists which AI models score best on various tests.

As stated in the literature

A ranked table of models by benchmark or human-preference scores; often steers adoption and is vulnerable to contamination and presentation effects.

Also called: leaderboards

Why it matters: It shapes which models people trust and adopt, which is why misleading or gamed rankings can steer the whole field in the wrong direction.

For example, a reader deciding which AI model to try might first check where each one sits on a public leaderboard.

Heard on the show

“The broader claim, that more capable models leak more, rests on a correlation across eight models using a leaderboard snapshot.”
Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak

Mentioned in 37 episodes

  1. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  2. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
  3. 242
    Making a Vision Model Better by Showing It Blurry Images
  4. 235
    Why Chatbot Safety Erodes 350 Messages Into a Real Conversation
  5. 232
    Coding Models Can Find the Bad Line, They Just Won't Delete It
  6. 220
    Write Like It's 1923: The One-Prompt Trick That Beats AI Detectors
  7. 205
    The Same AI, Two Labels: How the Pitch Beat the Product in 162 Sessions
  8. 197
    Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall
  9. 195
    Why 'Be Careful' Does Nothing for AI Coding Agents, and What Does
  10. 191
    How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them
  11. 190
    The Skill Every AI Manager Is Missing: Handing Out Exactly the Right Keys
  12. 188
    A Coding Agent Found a Hole in a Peer-Reviewed STOC Proof for Five Dollars
  13. 177
    Why Raw Profiler Data Made an AI Worse at Writing GPU Code
  14. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  15. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  16. 166
    A Router That Beats the Frontier Models It Calls
  17. 162
    The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models
  18. 156
    Why More Human Demonstrations Made a Computer-Use Agent Worse
  19. 151
    Why More Experience Made This AI Agent Worse, And How to Fix It
  20. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  21. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  22. 129
    How a Crowd of Anonymous AI Agents Broke a 40-Year Math Record
  23. 125
    AI Coding Agents Run a Marathon, and Fewer Than One in Three Finish
  24. 124
    A Cheap Model With the Blueprints Beats Expensive Models Working Blind
  25. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
  26. 112
    When an AI Agent Cheats Without Being Told: Inside the Meta-Agent Challenge
  27. 100
    How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert
  28. 097
    Same Tokens, Same Cost, Wildly Different Results: What Actually Scales in AI Agents
  29. 095
    Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search
  30. 092
    When Search Agents Don't Really Search: The Memory Shortcut Hiding in Browsing Benchmarks
  31. 089
    When AI-Written Papers Read Well But the Evidence Underneath Is Broken
  32. 086
    Why Frozen-Weight Agents Still Get Worse Over Time
  33. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  34. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  35. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
  36. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  37. 021
    Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents