Glossary · Term

tau-bench

← all terms

Definition

Plain language

A benchmark that tests AI agents on realistic customer-service phone-call style conversations.

As stated in the literature

A multi-turn agentic benchmark covering Retail, Airline, and other domains, evaluated with pass@k reliability metrics; distinct from tau2-bench, which extends it with additional tool environments.

Also called: τ-bench, tau bench

Why it matters: It exposes how reliably agents handle real customer-service workflows, where one wrong step can violate policy or anger a user.

For example, an agent must handle a multi-turn airline-rebooking call, looking up the customer's reservation and applying the right fare rules.

Heard on the show

“Some of the benchmarks here — tau-bench and tau2-bench in particular — report scores under what's called pass-cubed, where a task only counts as solved if the agent succeeds on three independent runs.”
Episode 071 — When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface

Related terms