Definition
Plain language
A benchmark where AI agents have to complete tasks on simulated websites.
As stated in the literature
A benchmark for web-based autonomous agents, providing realistic simulated websites and multi-step tasks for end-to-end evaluation.
Also called: WebArena-Lite
Why it matters: It provides a controlled, reproducible way to evaluate web agents that doesn't break every time real websites change their layout.
For example, an agent might be asked to find the cheapest item in a category on a fake shopping site and add it to a cart.
Heard on the show
“Task 429 on a benchmark called WebArena-Lite.”Episode 008 — Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps