Definition
Data poisoning is an attack that corrupts a model's training or retrieval data so that its behavior is subtly manipulated at inference time. In systems like Vis-Poison, a single carefully edited image inserted into a retrieval corpus can be enough to mislead every downstream answer that draws on it, showing how small, targeted corruptions can undermine trust in retrieval-augmented pipelines far out of proportion to their size.
Episodes covering this
Worth reading next
Papers we haven't done a deep dive on yet, but would recommend on this topic.
- TruFor: Leveraging All-Round Clues for Trustworthy Image Forgery Detection and Localization
- PoisonedRAG: Knowledge Poisoning Attacks to Retrieval-Augmented Generation of Large Language Models
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training