Definition
Plain language
A benchmark of real web tasks used to evaluate browsing agents.
As stated in the literature
A web-agent evaluation suite covering hundreds of tasks across real-world websites, used as a standard reference for browsing-agent generalization.
Why it matters: It provides a common yardstick for whether web agents actually generalize to the messy, varied websites people use.
For example, a browsing agent might be asked to find a one-way flight from Boston to Tokyo on a real airline site as one of the benchmark tasks.
Heard on the show
“They don't run it on alternative benchmarks like WebGames or Mind2Web either, which would have been useful generalization checks.”Episode 008 — Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps