Definition
Plain language
An attack that plants one subtly edited real photo so an AI search system gives the wrong answer.
As stated in the literature
A black-box, text-free multimodal RAG poisoning framework using a planner-editor-verifier loop to apply a minimal, query-relevant edit to a genuine image, preserving retrieval embedding proximity while flipping the answer.
Also called: VisPoison
Why it matters: Because the attack carries no text and barely changes the image, it defeats both content filters and the usual forgery detectors at once.
For example, a real photo of a landmark is altered in one small spot so a search system still retrieves it normally but the model reads the wrong detail off it.
Heard on the show
“The paper is called Vis-Poison, out of a team led by Rujin Liang, and the threat model is about as stripped-down as these things get.”Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It