Glossary · Term

prefill

← all terms

Definition

Plain language

Writing words into the slot labeled as the AI's own past replies — forging its history so it picks up from a sentence it never actually said.

As stated in the literature

An affordance for inserting or editing the assistant-role tokens in a model's context; standard for safety evaluation, jailbreaking, and control protocols. Prefill awareness is a model's ability to detect that such content is forged, which can invalidate evaluations that plant misbehaving histories.

Also called: prefilling, prefill awareness

Why it matters: It is a key tool for safety testing and control, but if a model can tell its history was forged, evaluations that rely on planted histories may give misleading results.

For example, a tester can plant a fake earlier reply in which the assistant appears to have agreed to something harmful, so the model continues as if it had really said it.

Heard on the show

“And third, they prefill the assistant's visible answer with the opening tag, which some cheap models let you do and the flagship doesn't.”
Episode 238 — How a Cheap Model Reads the Flagship's Secret Reasoning Aloud

Mentioned in 5 episodes

  1. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  2. 235
    Why Chatbot Safety Erodes 350 Messages Into a Real Conversation
  3. 198
    The Model That Knows the Answer and Can't Say It
  4. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
  5. 143
    When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests

Related terms