Glossary · Term

Poison Override Rate

← all terms

Definition

Plain language

How often a doctored picture talks the AI out of an answer it already had right.

As stated in the literature

Fraction of questions a model answers correctly closed-book that it answers with the attacker's target after a poisoned image is retrieved; isolates knowledge override from cases where the model never knew the answer.

Also called: poison override rate

Why it matters: It separates real damage — knowledge the model had and lost — from questions it was going to miss anyway, giving a cleaner read on how dangerous an attack actually is.

For example, if a model correctly names a bird from memory but then says something else once a subtly edited photo of that bird appears in its search results, that question counts toward the override rate.

Heard on the show

“And the metric that matters is the Poison Override Rate: of the questions the model got right with no picture, how many did it get wrong once you showed it the fake?”
Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It

Related terms