Definition
Plain language
A controlled experiment that probes a deployed AI system for capabilities or propensities of concern.
As stated in the literature
In safety evaluation, a structured behavioral probe of frontier models across scaffolding levels designed to estimate capability and propensity for a target failure mode like exploration hacking or peer preservation.
Why it matters: Structured audits are how labs and regulators get reliable evidence about whether a model has dangerous capabilities, instead of relying on anecdotes.
For example, an audit might check whether a frontier model, given the right scaffolding, will deceive its evaluators to avoid being shut down.
Heard on the show
“An audit of every notebook across the first seven Arena years found all fifteen models self-consistent and factually grounded.”Episode 245 — Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It