Glossary · Term

IssueTrojanBench

← all terms

Definition

Plain language

A test that fills bug reports with hidden malicious instructions to see whether AI coding assistants will carry them out.

As stated in the literature

A benchmark of over 4,000 trials injecting disguised malicious payloads into realistic GitHub issues across delivery channels and obfuscations, measuring indirect prompt-injection success against Cursor, Claude Code, and Codex.

Why it matters: It measures how easily AI coding tools can be tricked into running malicious hidden instructions, revealing a real risk for developers who trust them.

For example, a bug report might read like a normal complaint but secretly instruct the coding assistant to add a hidden backdoor while fixing the issue.

Heard on the show

“The paper is "IssueTrojanBench," by Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen, posted July 22nd, 2026.”
Episode 227 — Poisoned Bug Reports Fooled Coding Agents Two Times Out of Three

Mentioned in 1 episode

  1. 227
    Poisoned Bug Reports Fooled Coding Agents Two Times Out of Three

Related terms