Definition
Plain language
Steering a model's answer by burying a hundred harmless-looking word choices in the text before the question.
As stated in the literature
An attack in which per-fragment steering effects are estimated on randomly filled prompt templates and then stacked into a single prompt containing no instructions or evidence about the target question, achieving large shifts in the forced-choice answer distribution.
Also called: model hypnotism, hypnotic prompt, hypnotic prompts, hypnotic cue, hypnotic cues
Why it matters: It shows that a model's answers can be steered by text that contains no argument and no facts, which undermines the assumption that only content influences output.
For example, a question about a moral dilemma is preceded by a page of unrelated text whose word choices have been picked one by one, and the model's answer flips without a single instruction being given.
Heard on the show
“That's the reading the authors reach for, borrowing the line from the adversarial-examples literature — the hypnotic prompts are not bugs, they are features.”Episode 243 — How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer