Definition
Plain language
A benchmark that tests AI agents on realistic customer-service phone-call style conversations.
As stated in the literature
A multi-turn agentic benchmark covering Retail, Airline, and other domains, evaluated with pass@k reliability metrics; distinct from tau2-bench, which extends it with additional tool environments.
Also called: τ-bench, tau bench
Why it matters: It exposes how reliably agents handle real customer-service workflows, where one wrong step can violate policy or anger a user.
For example, an agent must handle a multi-turn airline-rebooking call, looking up the customer's reservation and applying the right fare rules.
Heard on the show
“Some of the benchmarks here — tau-bench and tau2-bench in particular — report scores under what's called pass-cubed, where a task only counts as solved if the agent succeeds on three independent runs.”Episode 071 — When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface