Definition
Plain language
A set of examples used to check progress while you're still improving a model.
As stated in the literature
Split used for model selection and hyperparameter tuning; provides the feedback signal distinct from the final test set.
Also called: validation set, validation scores
Why it matters: It gives you a feedback signal for making choices without burning through the final test set, which must stay untouched to remain trustworthy.
For example, after each tweak to a model's settings you check its accuracy on these held-aside examples to decide whether the tweak helped before trying the next one.
Heard on the show
“It runs MCTS — Monte Carlo Tree Search, the same family of algorithms that powered AlphaGo — over thousands of possible workflow variants, scoring each on a small validation set, keeping the best one.”Episode 013 — Why Search Keeps Rediscovering the Same Workflow, and What That Means