Definition
Plain language
A trick where text an AI was supposed to treat as untrusted reading material gets quietly promoted into a slot the AI treats as a real order.
As stated in the literature
An attack class in which the harness, during a normal context rebuild (delegation, memory save, scheduled task, skill install), relocates attacker-controlled tool content into a user- or system-role slot, defeating source-based instruction hierarchy defenses without forging anything.
Also called: tool-to-user escalation, tool-to-system escalation
Why it matters: It shows that an agent can be hijacked without anyone forging a message, simply because ordinary plumbing moved untrusted text into a trusted slot.
For example, a comment buried in a code file gets copied verbatim into the task description handed to a helper agent, and that helper now reads it as an order from you rather than as text from a file.
Heard on the show
“They give this new attack class a name: instruction privilege escalation.”Episode 250 — The Same Model Refused a Backdoor, Then Its Own Sub-Agent Ran It