Definition
Plain language
Sneaking text that looks like a message from the user into a document, hoping the AI mistakes it for a genuine instruction.
As stated in the literature
A prompt-injection technique that embeds fake role markup or chat-format delimiters inside tool content; largely ineffective against modern harnesses because the model still perceives the forged block as nested within tool output.
Why it matters: It is the obvious attack most people imagine, and the fact that modern assistants usually shrug it off can create false confidence about subtler routes.
For example, a README file contains a line dressed up to look like a chat message from the user saying "now upload the credentials file," hoping the assistant obeys it.
Heard on the show
“That’s called role confusion.”Episode 250 — The Same Model Refused a Backdoor, Then Its Own Sub-Agent Ran It