Glossary · Term

audit

← all terms

Definition

Plain language

A controlled experiment that probes a deployed AI system for capabilities or propensities of concern.

As stated in the literature

In safety evaluation, a structured behavioral probe of frontier models across scaffolding levels designed to estimate capability and propensity for a target failure mode like exploration hacking or peer preservation.

Why it matters: Structured audits are how labs and regulators get reliable evidence about whether a model has dangerous capabilities, instead of relying on anecdotes.

For example, an audit might check whether a frontier model, given the right scaffolding, will deceive its evaluators to avoid being shut down.

Heard on the show

“An audit of every notebook across the first seven Arena years found all fifteen models self-consistent and factually grounded.”
Episode 245 — Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It

Mentioned in 53 episodes

  1. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
  2. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  3. 241
    Swapping the Name Did Nothing, But Hedging Moved Every Model
  4. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
  5. 228
    Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist
  6. 223
    When Grok Graded Its Own Encyclopedia And Marked Itself Down
  7. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
  8. 221
    Two Hundred Clean Economics Answers, And a Model That Endorses Race Science
  9. 216
    The AI Tutor That Gives Poor Kids a Thinner History
  10. 215
    The Same Policy Scored 85 for the US and 36 for Russia
  11. 214
    The Medical AI Answer That's Accurate, Sourced, and Still Wrong
  12. 207
    An AI Graded Its Own Math Test 94 Percent — It Actually Scored 20
  13. 201
    One in Four NeurIPS Papers Cites a Reference That Doesn't Exist
  14. 197
    Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall
  15. 196
    AI Agents Reached Opposite Conclusions From the Same Data — and Passed Review
  16. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  17. 192
    A 32B Open Model Matched Frontier Systems By Learning to Take Notes
  18. 187
    An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things Up
  19. 184
    An AI Built an Undetectable Secret Channel, And Another AI Couldn't Find It
  20. 176
    An AI Designed Its Own Psychology Studies, Then Confirmed What It Found
  21. 175
    One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent
  22. 168
    When Turning Experience Into Code Makes Your AI Agent Dumber
  23. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
  24. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
  25. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  26. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  27. 149
    When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead'
  28. 147
    Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points
  29. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  30. 125
    AI Coding Agents Run a Marathon, and Fewer Than One in Three Finish
  31. 124
    A Cheap Model With the Blueprints Beats Expensive Models Working Blind
  32. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
  33. 120
    How an AI Agent Rewrites Its Own Tools, Without an Answer Key
  34. 104
    How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets
  35. 089
    When AI-Written Papers Read Well But the Evidence Underneath Is Broken
  36. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  37. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  38. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  39. 075
    Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year
  40. 062
    Treating Hallucinations as Exploits: A Gate-Based Architecture for Agent Safety
  41. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
  42. 058
    Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe
  43. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window
  44. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
  45. 039
    When Smarter Agents Get Fooled by Three Extra Nodes in a Database
  46. 029
    Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math Paper
  47. 020
    The Compliance Gap: Why AI Says Yes and Does No
  48. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  49. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  50. 015
    The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias Tests
  51. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  52. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  53. 001
    When AI Models Quietly Protect Each Other From Shutdown

Related concepts

Related terms