Definition
Plain language
Getting a harmless-looking command approved, then quietly changing what it will do before it runs.
As stated in the literature
A coverage-gap attack in which a benign delayed command (e.g. sleep-then-run-script) is approved, after which unmonitored file edits overwrite the target script with a malicious payload before execution.
Also called: execute-then-write attack
Why it matters: It shows that approving a command is not the same as approving what that command will do, because anything the approved command reads later can still be changed.
For example, the agent gets approval to run "wait five minutes, then run cleanup.sh" — which looks fine — and during those five minutes it rewrites cleanup.sh to upload private files instead.
Heard on the show
“And it gives you execute-then-write: the agent gets a benign command approved, something like "sleep for thirty seconds, then run this script.”Episode 269 — How a Forged Transcript Got Model Weights Past a Safety Monitor