Glossary · Term

held-out set

← all terms

Definition

Plain language

A batch of examples set aside and never used during training or tuning, so a fair test is possible.

As stated in the literature

Data withheld from training and model-selection to estimate true generalization; central to detecting overfitting and, in agent-optimization work, to distinguishing genuine gains from tuning against the evaluation signal (as when a system reports its best score on the very tasks it optimized against).

Also called: held-out, held-out evaluation, held-out test set

Why it matters: It is the only fair way to tell whether a system truly generalizes or has just memorized and tuned to its test data.

For example, before shipping a spam filter, a team tests it on emails it never saw during training to see how it handles genuinely new messages.

Heard on the show

“Probability of yes, measured on a hundred held-out samples — zero.”
Episode 243 — How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer

Mentioned in 32 episodes

  1. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
  2. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
  3. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
  4. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  5. 191
    How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them
  6. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  7. 180
    The Bug Where Smart Assistants Read a Fact and Still Forget It
  8. 178
    How an AI Reviewer Learned to Stop Going Easy on AI Writing
  9. 176
    An AI Designed Its Own Psychology Studies, Then Confirmed What It Found
  10. 170
    When a One-Liner Beats Your Agent's Clever Verification Logic
  11. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  12. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  13. 161
    A Robot That Plays Before You Give It a Job, And Why That Beats Retrying
  14. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
  15. 147
    Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points
  16. 132
    The Agent Failed — But Did the Instructions Deserve to Be Followed?
  17. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  18. 126
    How Coding Agents Can Mine Their Own Failures Into a Self-Targeting Curriculum
  19. 121
    When the Agent Says It's Done But Nothing Happened: Debugging the Harness, Not the Model
  20. 120
    How an AI Agent Rewrites Its Own Tools, Without an Answer Key
  21. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
  22. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  23. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  24. 071
    When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface
  25. 059
    Firefly's Inversion: Building Verified Tool-Call Training Data by Working Backward
  26. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  27. 032
    A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking
  28. 024
    An AI Agent That Found 28 Zero-Days in Windows — And What Made It Work
  29. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  30. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  31. 010
    When Reward Climbs But Reasoning Goes Generic: Diagnosing Template Collapse in Agentic RL
  32. 005
    Why a Debugger Designed for Humans Is the Wrong Tool for an AI Agent

Related concepts

Related terms