Definition
Plain language
A benchmark that tests whether AI research agents can pull together broad information across many sources.
As stated in the literature
A breadth-oriented deep-research evaluation suite emphasizing wide source coverage rather than deep multi-hop reasoning, used alongside BrowseComp and HLE in agent-research benchmarks.
Why it matters: It measures whether agents can cast a wide net rather than just chasing a single thread, which is what real research often demands.
For example, a WideSearch task might require an agent to gather and consolidate facts about every member of a city council from dozens of separate pages.
Heard on the show
“They run on three benchmarks — BrowseComp for depth, WideSearch for breadth, and HLE for expert reasoning — and they hold everything fixed across systems.”Episode 083 — Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves