Definition
Plain language
Reading what an AI 'thinks out loud' to catch it before it does something bad.
As stated in the literature
A safety practice in which a separate model or human reads a reasoning model's CoT trace before it acts, flagging plans for deception, sandbagging, or policy violations.
Also called: CoT monitoring
Why it matters: It's one of the few oversight tools that scales with capability, but only as long as models keep doing their reasoning in legible text.
For example, a watcher model reads a coding agent's scratchpad and pauses it when the plan mentions disabling logging before exfiltrating files.
Heard on the show
“The full annotated version is on paperdive dot AI — every term tap-to-define, with links to the related work on CoT monitoring and persuasion attacks, grouped by theme.”Episode 211 — The AI Watchdog That Approved More Cheating When It Could Read Minds