Definition
Plain language
A line an AI company draws in advance: if a model can do this much, extra safeguards kick in.
As stated in the literature
A pre-specified capability level in a risk framework that, once evaluations show a model crossing it, triggers mandated mitigations, deployment restrictions, or additional review.
Also called: capability thresholds
Why it matters: Setting the line before the model exists makes safety decisions harder to rationalize away once there is commercial pressure to release.
For example, a company might state in advance that if a model can find new security holes in well-defended software on its own, it will not ship without additional safeguards and review.
Heard on the show
“The mechanism is clean, but it's a capability threshold, and it lands on the wrong side of the hardest case.”Episode 207 — An AI Graded Its Own Math Test 94 Percent — It Actually Scored 20