Definition
Plain language
When the way you pick cases skews what you find, even if each case is studied honestly.
As stated in the literature
A statistical pattern where the sample selection process correlates with the outcome variable, biasing estimates; flagged for instance in QUEST's auto-formalization filtering and SIA's verifier-friendly benchmark choices.
Why it matters: It's the silent enemy of benchmark-driven research: the very filter that lets you build a dataset cheaply can guarantee the data isn't representative.
For example, if you only train an AI on math problems whose answers your auto-grader can verify, the resulting model gets very good at easily-checkable math — and possibly worse elsewhere.
Heard on the show
“Exploration bias, then selection bias.”Episode 196 — AI Agents Reached Opposite Conclusions From the Same Data — and Passed Review