Glossary · Term

held-out set

← all terms

Definition

Plain language

A batch of examples set aside and never used during training or tuning, so a fair test is possible.

As stated in the literature

Data withheld from training and model-selection to estimate true generalization; central to detecting overfitting and, in agent-optimization work, to distinguishing genuine gains from tuning against the evaluation signal (as when a system reports its best score on the very tasks it optimized against).

Also called: held-out, held-out evaluation, held-out test set

Why it matters: It is the only fair way to tell whether a system truly generalizes or has just memorized and tuned to its test data.

For example, before shipping a spam filter, a team tests it on emails it never saw during training to see how it handles genuinely new messages.

Heard on the show

“Each task is a real GitHub issue from a popular Python library, with held-out tests.”
Episode 005 — Why a Debugger Designed for Humans Is the Wrong Tool for an AI Agent

Related concepts

Related terms