Definition
Plain language
Data that looks like the examples a system was trained on, so it handles it reliably.
As stated in the literature
Inputs drawn from the same distribution as a model's training data, where its calibration and accuracy hold; contrasted with out-of-distribution inputs that fall in blind spots.
Also called: in distribution
Why it matters: Systems tend to be accurate and confident on in-distribution data but stumble on anything unfamiliar, so knowing the boundary tells you where to trust the output.
For example, a model trained only on photos of cats and dogs handles a new cat photo reliably because it falls within what it has already seen.
Heard on the show
“That truncated pipeline still beats every mixed-policy method on the in-distribution math benchmarks.”Episode 009 — How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers