Definition
Plain language
OpenAI Codex's built-in reviewer that checks a coding assistant's actions before they run.
As stated in the literature
The automated action-approval monitor shipped with Codex; in this paper it is reconstructed as an approximation (with sandboxing deliberately excluded) to measure classifier-level attack success.
Why it matters: It is one of the main real-world examples of shipping a model-based approver instead of asking users to confirm every step, so how well it holds up is a practical question rather than an academic one.
For example, when the coding assistant proposes running a command that sends a file to an outside address, this reviewer inspects the proposed action first and can refuse it before anything runs.
Heard on the show
“Claude Code ships this as Auto Mode, and Codex ships something similar called Guardian.”Episode 269 — How a Forged Transcript Got Model Weights Past a Safety Monitor