Definition
Plain language
The spread of answers a model tends to produce when you ask it the same thing many times.
As stated in the literature
The output distribution induced by a model and decoding settings; repeated sampling cannot recover set members that lie outside its support.
Why it matters: If an answer never appears in the model's output, no amount of retrying will find it, so repeated sampling fixes flakiness but not blind spots.
For example, asking the model for European capitals twenty times might surface Paris and Berlin every time and Valletta never, so Valletta is simply outside what it produces.
Heard on the show
“Sampling the set ten times and filtering with the model's own judgment made things worse at three of four scales, because the missing members simply aren't in the sampling distribution.”Episode 233 — Why a Model Can Grade an Answer But Not Write the Answer Key