All posts

Guides2 min read

τ-bench, explained: the benchmark behind our numbers

Every number on this site comes from τ-bench retail. What the benchmark is, how it grades an agent, why we use it, and what it can't tell you.

Every result on this site names its source: τ-bench retail, with simulated customers. Here is what that benchmark is, and why we chose it.

What τ-bench is

τ-bench is a public benchmark for customer-service agents, introduced in the paper "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (ICLR 2025). It has two domains, retail and airline. In each, an agent gets the domain's tools and a policy it has to follow, and a language model plays the customer.

  • Tools: the agent looks things up and changes things, like finding an order or issuing a refund.
  • Policy: rules the agent must follow, like what can be refunded and when.
  • A simulated customer: a language model with a goal, who talks to the agent the way a person would.

How it grades an agent

τ-bench ignores how well the agent writes. When the conversation ends, it compares the state of the database with the state the task should have produced. Either the right order was refunded, or it wasn't.

It also asks for reliability, not one lucky run. Its pass^k measure checks whether an agent gets the same task right across several tries.

Graded on what the agent did, not on what it said.

Why it's the right test for us

  • It is the work AgentCompile is built for: an agent with tools, a policy and a customer, doing the same jobs again and again.
  • It is graded on actions, so "same answers" is something we can measure, not something we claim.
  • It is public, so anyone can check the setup.

How we use it

  • Held out: AgentCompile learns only from training tasks and is measured on test tasks it never saw.
  • Same agent both ways: the agent model, the simulated customer, the prompts and the time window are the same with and without AgentCompile.
  • Every model call counted, AgentCompile's own included.

What it can't tell you

  • Our results are in one domain, retail.
  • The customers are simulated. None of it is customer data.
  • Your agent's jobs are its own. A pilot measures your own traffic.

See every result, with its setup and limits. Read the research

Join the beta. 10 spots.

If your company runs an AI agent in production, we'd like to compile its most repeated jobs with you.

Book a call

cal.com/agent-compile/beta

To start, we'll ask to see your agent's logs.