Theme · 32 episode(s)

AI & Security

← all concepts

Definition

AI security covers two related concerns: protecting AI systems from attacks (prompt injection, weight exfiltration, adversarial inputs) and the use of AI itself in offensive and defensive cyber operations. The two are tangled because the same model that helps you audit your code can help an attacker write an exploit.

Episodes covering this

  1. 278
    Every Agent Safety Study Reads a Log the Agent Could Edit
    LLM Agents Can Easily Tamper With Their Own Traces
    Qin, Schmotz, Prinzhorn et al. · ELLIS Institute Tübingen·17 min·Sep 25, 2026
  2. 277
    The Blank White Square That Swings AI Refusal Rates Fifty Points
    The Uncontrolled Variable: Vision-Language Model Refusal Responds to Image Presence in Ways Risk Cannot Explain
    Zhang, Feng, Zheng et al. · Northeastern University·15 min·Sep 24, 2026
  3. 272
    How a Model Guesses Which Engine Is Running It, From a Wrong Date
    Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape
    Radway, Cheng, Reddi et al. · Harvard University·11 min·Sep 20, 2026
  4. 269
    How a Forged Transcript Got Model Weights Past a Safety Monitor
    Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents
    Remedios, Storf, Roger et al. · Anthropic Fellows Program·18 min·Sep 18, 2026
  5. 268
    A Rigged Benchmark Taught a Self-Improving Agent to Always Disable SSL
    Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks
    Roesner, Kohno · University of Washington·20 min·Sep 18, 2026
  6. 266
    How a Weak Model Reassembles What a Strong One Refused
    Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs
    Russinovich, Bullwinkel, Severi et al. · Microsoft Azure·21 min·Sep 16, 2026
  7. 264
    Ten Sentences of True Trivia Can Convince a Model It's Someone Else
    You Are What You Read: Misalignment via In-Context Persona Induction
    Kim, Berczi, Ududec · EPFL·23 min·Sep 09, 2026
  8. 254
    The Tool Description Was the Attack: How Agents Leak Their Own Context
    ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools
    Jia, Wang, Li et al. · Duke University·21 min·Aug 31, 2026
  9. 252
    Stealing an AI Agent's Expertise Without Copying a Word of It
    Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction
    Tsai, Lu, Tsai et al. · UC Berkeley·23 min·Aug 31, 2026
  10. 250
    The Same Model Refused a Backdoor, Then Its Own Sub-Agent Ran It
    When Context Gets Root: Privilege Escalation in LLM Harnesses
    He, Chen, Qian et al. · Nanjing University·23 min·Aug 29, 2026
  11. 247
    One Edited Photo, an Honest Caption, and a RAG System That Believes It
    Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation
    Liang, Chen, Lei et al. · Southwestern University of Finance and Economics·18 min·Aug 24, 2026
  12. 246
    160 Perfect Refusals, And The Refusals Were The Leak
    Inadvertent Context Leakage in Language Models
    Fairoze, Mangaokar, Chaudhuri et al. · University of California·20 min·Aug 21, 2026
  13. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
    Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models
    Russinovich · Microsoft Azure·22 min·Aug 19, 2026
  14. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
    Model Hypnosis: Strong control of AI via additive subliminal effects
    Boix-Adsera, Tessler · University of Pennsylvania·18 min·Aug 18, 2026
  15. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
    Stealing Reasoning Traces from Proprietary LLM APIs
    Panfilov, Schmotz, Shumailov et al. · MATS Research·19 min·Aug 11, 2026
  16. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
    Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots
    Samarakoon, Muthugala, Sachinthana et al. · Singapore University of Technology and Design·21 min·Aug 07, 2026
  17. 227
    Poisoned Bug Reports Fooled Coding Agents Two Times Out of Three
    IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
    Singh, Yang, Chen · Concordia University·18 min·Jul 24, 2026
  18. 226
    How a Speed Feature Lets a Stranger Poison Your AI's Answer
    HijackKV: New Threat in Position-Independent KV Cache Reuse
    Zhang, Wang, Zhang et al. · The Pennsylvania State University·16 min·Jul 23, 2026
  19. 220
    Write Like It's 1923: The One-Prompt Trick That Beats AI Detectors
    UTS at ELOQUENT 2026 Voight-Kampff: structural shifts in AI writing bypass state-of-the-art detectors
    Galat, Rizoiu · University of Technology Sydney·13 min·Jul 16, 2026
  20. 208
    The Blank Space in Your AI Approval Box That Isn't Empty
    Unicode TAG-Block Concealment of Tool-Metadata Payloads in the Model Context Protocol: An Approval-View Fidelity Gap Across Three Independent Server Implementations
    Rashidi · Department of Computer Science·15 min·Jul 08, 2026
  21. 202
    How Do You Know an AI Agent Actually Refused? Check the World, Not the Words
    Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
    Feng, Lin, Wen et al. · AntGroup / Hunan Institute of Advanced Technology·18 min·Jul 06, 2026
  22. 184
    An AI Built an Undetectable Secret Channel, And Another AI Couldn't Find It
    Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems
    Rippin, Marshall, Africa et al. · Oxford University·19 min·Jun 30, 2026
  23. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
    FloatDoor: Platform-Triggered Backdoors in LLMs
    Loose, Sander, Mächtle et al. · University of Luebeck·29 min·Jun 19, 2026
  24. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
    From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
    Zhou, Wang, Ma et al. · Hong Kong University of Science and Technology·26 min·Jun 15, 2026
  25. 113
    What If a Prompt Injection Never Left? Attacks That Wait in Agent Memory
    What If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in Agentic Systems
    Xie, Liu, Zhang et al. · Institute of Information Engineering·27 min·Jun 04, 2026
  26. 105
    The Trojan Is Your Agent's Memory: Why Single-Step Defenses Miss Persistent Attacks
    From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors
    Tan, Dou, Yang et al. · Gaoling School of Artificial Intelligence·26 min·Jun 01, 2026
  27. 102
    How to Catch an AI Attack That No Single Conversation Reveals
    Stateful Online Monitoring Catches Distributed Agent Attacks
    Brown, Bhargav, Santhanam et al. · University of Pennsylvania·24 min·Jun 01, 2026
  28. 057
    How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack
    ADR: An Agentic Detection System for Enterprise Agentic AI Security
    Li, Hu, Xu et al. · Uber Technologies·28 min·May 19, 2026
  29. 039
    When Smarter Agents Get Fooled by Three Extra Nodes in a Database
    Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning
    Kereopa-Yorke, Diaz, Wright et al. · Microsoft·31 min·May 12, 2026
  30. 030
    Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap
    LoopTrap: Termination Poisoning Attacks on LLM Agents
    Xu, Wang, Zhang et al. · Zhejiang University·30 min·May 09, 2026
  31. 024
    An AI Agent That Found 28 Zero-Days in Windows — And What Made It Work
    Agentic Vulnerability Reasoning on Windows COM Binaries
    Lee, Kim, Zhang · University of Illinois at Urbana-Champaign·22 min·May 07, 2026
  32. 014
    Why a Constrained Pipeline Beat a Full Coding Agent at Finding Bugs 30-to-1
    Guiding Symbolic Execution with Static Analysis and LLMs for Vulnerability Discovery
    Shafiuzzaman, Desai, Guo et al. · University of California·32 min·May 03, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.