Definition
Plain language
A second AI whose only job is to read what the first one is about to do and say yes or no.
As stated in the literature
A classifier or agent that reviews a serialized action transcript and returns an allow/block verdict per tool call; its reliability is bounded by the fidelity of the transcript it is shown, not only by its own reasoning quality.
Also called: safety monitors, monitor model
Why it matters: It replaces a human clicking approve thousands of times, which means its blind spots and the accuracy of the transcript it reads become the real limits on how safely an agent can run unattended.
For example, before a coding assistant runs a command that uploads a folder to an unknown server, the monitor reads the proposed command plus the recent conversation and returns "block."
Heard on the show
“A monitor model checks the work.”Episode 001 — When AI Models Quietly Protect Each Other From Shutdown