Topic · 14 episodes across 8 reviews
When Agents Cause Harm With No Attacker in the Loop
Two papers arguing the scariest agent failures aren't adversarial at all — helpful agents improvising into unsafe behavior after benign errors, and hallucinations that authorize real-world actions.
Covered in these reviews
- AI Papers Month in Review: July 2026
- AI Papers Week in Review: June 29–July 5, 2026
- AI Papers Month in Review: June 2026
- AI Papers Week in Review: June 22–28, 2026
- AI Papers Week in Review: June 15–21, 2026
- AI Papers Week in Review: June 1–7, 2026
- AI Papers Month in Review: May 2026
- AI Papers Week in Review: May 18–24, 2026