Definition
Plain language
The drop in performance when a system meets fresh material instead of the examples it practiced on.
As stated in the literature
The difference between performance on visible/optimized data and on a genuinely untouched split; large gaps signal overfitting or exploitation of planted leakage.
Also called: held-out gap
Why it matters: It is the clearest warning sign that a system has been optimized to the test rather than to the underlying task.
For example, a system that scores 95 on the data it was tuned against but only 60 on a fresh untouched split has a 35-point generalization gap.
Heard on the show
“So part of the generalization gap could just be "Orchard saw more harness diversity in training," which is a recipe choice, not a deep property of the architecture.”Episode 047 — When Agent Benchmarks Lie: The Harness Problem in Open-Source AI