Definition
Plain language
A benchmark of enterprise-workflow tasks for AI agents.
As stated in the literature
An evaluation suite of business-style multi-step tasks spanning analysis, reporting, and communications, used to stress-test general LLM agents.
Also called: The Agent Company
Why it matters: It evaluates whether general LLM agents can actually carry out the kinds of office tasks vendors keep promising they'll automate.
For example, an agent might be asked to read a quarterly sales spreadsheet, write a memo summarizing trends, and email it to the right team.
Heard on the show
“The Agent Company is five apps, a hundred-seventy-five tasks.”Episode 017 — When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers