Concept · 1 episode(s)

Data Poisoning

← all concepts

Definition

Data poisoning is an attack that corrupts a model's training or retrieval data so that its behavior is subtly manipulated at inference time. In systems like Vis-Poison, a single carefully edited image inserted into a retrieval corpus can be enough to mislead every downstream answer that draws on it, showing how small, targeted corruptions can undermine trust in retrieval-augmented pipelines far out of proportion to their size.

Episodes covering this

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.