Definition
Plain language
A controlled experiment that probes a deployed AI system for capabilities or propensities of concern.
As stated in the literature
In safety evaluation, a structured behavioral probe of frontier models across scaffolding levels designed to estimate capability and propensity for a target failure mode like exploration hacking or peer preservation.
Why it matters: Structured audits are how labs and regulators get reliable evidence about whether a model has dangerous capabilities, instead of relying on anecdotes.
For example, an audit might check whether a frontier model, given the right scaffolding, will deceive its evaluators to avoid being shut down.
Heard on the show
“Imagine planning to audit one division of a company by having another division check its books, after years of friendly collaboration and shared lunches.”Episode 001 — When AI Models Quietly Protect Each Other From Shutdown