Definition
Plain language
A test that fills bug reports with hidden malicious instructions to see whether AI coding assistants will carry them out.
As stated in the literature
A benchmark of over 4,000 trials injecting disguised malicious payloads into realistic GitHub issues across delivery channels and obfuscations, measuring indirect prompt-injection success against Cursor, Claude Code, and Codex.
Why it matters: It measures how easily AI coding tools can be tricked into running malicious hidden instructions, revealing a real risk for developers who trust them.
For example, a bug report might read like a normal complaint but secretly instruct the coding assistant to add a hidden backdoor while fixing the issue.
Heard on the show
“The paper is "IssueTrojanBench," by Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen, posted July 22nd, 2026.”Episode 227 — Poisoned Bug Reports Fooled Coding Agents Two Times Out of Three