Definition
Plain language
The problem of supervising an AI that may be smarter than its supervisors.
As stated in the literature
The challenge of constructing reliable supervision for capable AI systems using weaker or cheaper overseers, including debate, prover-verifier games, and conservative-agency approaches.
Why it matters: As AI grows more capable than its supervisors, finding ways to keep it honest and on track becomes essential to deploying it safely.
For example, it asks how a person could reliably check the work of an AI that reasons faster and deeper than they can.
Heard on the show
“And from a safety standpoint, Finn, both are problems for scalable oversight, but they're problems with very different shapes.”Episode 001 — When AI Models Quietly Protect Each Other From Shutdown