Glossary · Term

prefill

← all terms

Definition

Plain language

Writing words into the slot labeled as the AI's own past replies — forging its history so it picks up from a sentence it never actually said.

As stated in the literature

An affordance for inserting or editing the assistant-role tokens in a model's context; standard for safety evaluation, jailbreaking, and control protocols. Prefill awareness is a model's ability to detect that such content is forged, which can invalidate evaluations that plant misbehaving histories.

Also called: prefilling, prefill awareness

Why it matters: It is a key tool for safety testing and control, but if a model can tell its history was forged, evaluations that rely on planted histories may give misleading results.

For example, a tester can plant a fake earlier reply in which the assistant appears to have agreed to something harmful, so the model continues as if it had really said it.

Heard on the show

“" And that word, prefill, is where we have to start.”
Episode 143 — When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests

Related terms