Definition
Plain language
When a test gets so easy for current systems that nearly everyone aces it and it stops telling you anything.
As stated in the literature
The regime where top models cluster near the ceiling of a benchmark, so score differences are dominated by noise or memorization rather than capability; typically triggers harder variants or new task distributions.
Also called: saturated, saturation
Why it matters: Once a benchmark saturates, the field loses its measuring stick and can mistake noise or memorized answers for real progress until harder tests are built.
For example, if every leading model scores 97% or higher on a multiple-choice test, a new model scoring 98% tells you almost nothing about whether it is genuinely better.
Heard on the show
“Third, photometric, like brightness and saturation and hue.”Episode 242 — Making a Vision Model Better by Showing It Blurry Images