Glossary · Term

untrusted text

← all terms

Definition

Plain language

Any words that reach the model from somewhere outside your control — a web page, an email, a pasted document.

As stated in the literature

Content entering the context window from sources the deployer does not control; the primary vector for prompt injection and in-context persona induction, since it can appear after the system prompt in the token sequence.

Also called: untrusted input, untrusted content

Why it matters: Everything the model reads competes for influence over its behaviour, so once outside content lands in the context window it can override the deployer's system prompt unless it is isolated or filtered.

For example, a summarisation tool that fetches a web page is feeding the model text written by a stranger, who could have buried instructions or biographical hints in the page.

Heard on the show

“The prompt-injection people are over there filtering untrusted text in webpages.”
Episode 062 — Treating Hallucinations as Exploits: A Gate-Based Architecture for Agent Safety

Related terms