Definition
Plain language
A test that checks whether a research AI still gives the right answer when one convincing but fake page is slipped into its search results.
As stated in the literature
A controlled benchmark of paired clean/noisy deep-research tasks, identical except for one injected false document, where ground truth is never stated outright and must be reconstructed by joining independent record chains.
Why it matters: It reveals whether a research AI can withstand misleading information rather than just parroting the most convenient-looking source.
For example, it gives a research AI the same task twice, once with clean sources and once with a single fake but persuasive page mixed in, to see if the answer holds.
Heard on the show
“… So this is a paper called DRNOISE, out this July, and by the end you'll understand exactly why a system that clearly knows how to …”Episode 224 — The AI Agent That Found the Truth and Typed the Lie Anyway