Glossary · Term

chain-of-thought monitoring

← all terms

Definition

Plain language

Reading what an AI 'thinks out loud' to catch it before it does something bad.

As stated in the literature

A safety practice in which a separate model or human reads a reasoning model's CoT trace before it acts, flagging plans for deception, sandbagging, or policy violations.

Also called: CoT monitoring

Why it matters: It's one of the few oversight tools that scales with capability, but only as long as models keep doing their reasoning in legible text.

For example, a watcher model reads a coding agent's scratchpad and pauses it when the plan mentions disabling logging before exfiltrating files.

Heard on the show

“You generate a big pile of plausible-looking documents about some topic — in this case, CoT monitoring.”
Episode 054 — When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window

Related terms