Glossary · Term

chain-of-thought monitoring

← all terms

Definition

Plain language

Reading what an AI 'thinks out loud' to catch it before it does something bad.

As stated in the literature

A safety practice in which a separate model or human reads a reasoning model's CoT trace before it acts, flagging plans for deception, sandbagging, or policy violations.

Also called: CoT monitoring

Why it matters: It's one of the few oversight tools that scales with capability, but only as long as models keep doing their reasoning in legible text.

For example, a watcher model reads a coding agent's scratchpad and pauses it when the plan mentions disabling logging before exfiltrating files.

Heard on the show

“The full annotated version is on paperdive dot AI — every term tap-to-define, with links to the related work on CoT monitoring and persuasion attacks, grouped by theme.”
Episode 211 — The AI Watchdog That Approved More Cheating When It Could Read Minds

Mentioned in 4 episodes

  1. 211
    The AI Watchdog That Approved More Cheating When It Could Read Minds
  2. 128
    How a Model Can Earn Full Reward and Still Resist Training
  3. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  4. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window

Related terms