Glossary · Term

chance level

← all terms

Definition

Plain language

The score you'd get by guessing at random.

As stated in the literature

The baseline accuracy of an uninformative predictor, e.g. 33% for a uniform three-way classification; probe results near chance indicate the target information is not linearly recoverable.

Also called: chance

Why it matters: Without knowing the chance level, a score can look impressive when it is actually indistinguishable from random guessing.

For example, on a quiz where every question has four options, a system that knows nothing at all will still get about 25% right, so 25% is the number any real result has to beat.

Heard on the show

“It never gets a chance to lie.”
Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It

Mentioned in 46 episodes

  1. 247
    One Edited Photo, an Honest Caption, and a RAG System That Believes It
  2. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  3. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
  4. 234
    Two Copies of Gemini Cooperated in a Game Where Betrayal Always Pays
  5. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
  6. 212
    The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know
  7. 210
    Same Website Request, Different Code — The Bias You Can't See
  8. 205
    The Same AI, Two Labels: How the Pitch Beat the Product in 162 Sessions
  9. 201
    One in Four NeurIPS Papers Cites a Reference That Doesn't Exist
  10. 196
    AI Agents Reached Opposite Conclusions From the Same Data — and Passed Review
  11. 184
    An AI Built an Undetectable Secret Channel, And Another AI Couldn't Find It
  12. 182
    How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%
  13. 179
    How DeepSeek Made One User Faster Without Slowing Down the Crowd
  14. 176
    An AI Designed Its Own Psychology Studies, Then Confirmed What It Found
  15. 175
    One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent
  16. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
  17. 172
    One Bad Token Can Sink a Model's Math, And You Can Delete It
  18. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  19. 162
    The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models
  20. 153
    Catching a Lie From the Inside, When the Words Look Completely Honest
  21. 143
    When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests
  22. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  23. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
  24. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  25. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  26. 130
    Why AI Agents Coordinate Better Through a Shared Board Than a Boss
  27. 128
    How a Model Can Earn Full Reward and Still Resist Training
  28. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  29. 113
    What If a Prompt Injection Never Left? Attacks That Wait in Agent Memory
  30. 107
    How a Market of Crippled AI Agents Outscored One Unrestricted Model
  31. 105
    The Trojan Is Your Agent's Memory: Why Single-Step Defenses Miss Persistent Attacks
  32. 100
    How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert
  33. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  34. 085
    Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction
  35. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  36. 070
    When Models Know the Answer But Say the Wrong Thing Anyway
  37. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  38. 063
    Why Web Agents Are Slow: A Compiler-Style Fix for Computer-Use Latency
  39. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  40. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  41. 033
    Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval
  42. 032
    A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking
  43. 020
    The Compliance Gap: Why AI Says Yes and Does No
  44. 018
    Language Models Compute the Rational Move, Then Override It
  45. 010
    When Reward Climbs But Reasoning Goes Generic: Diagnosing Template Collapse in Agentic RL
  46. 004
    The Sycophancy Circuit That Survives Alignment Training

Related terms