All research

Method1 min read

How we measure

Held out, same agent both ways, every model call counted, AgentCompile's own included. The rules every number on this site follows.

Key result

5,000+ benchmark conversations.

3 agent models: Gemini 2.5 Flash, Gemini 2.5 Pro, Claude Sonnet 4.5 · 808 automated tests

Setup

Benchmark
τ-bench, the public customer-service agent benchmark, retail domain
Held out
Learned only from training tasks, measured on test tasks it never saw
Same both ways
Agent model, simulated customer, prompts and time window
Counted
Every model call, AgentCompile's own included
Reported
Sample sizes and 95% intervals

The rules

  • Held out. AgentCompile learns only from training tasks and is measured on test tasks it never saw.
  • Same agent both ways. The agent model, the simulated customer, the prompts and the time window are the same with and without AgentCompile.
  • Every model call counted, AgentCompile's own included.
  • Sample sizes and 95% intervals, with every number.
  • Per-task results are available. Ask us.

How to read our numbers

  • Every number keeps its source: τ-bench retail, with simulated customers.
  • None of it is customer data. A pilot measures your own traffic.
  • Fewer agent calls is a result on its own. Don't turn it into a cost or a latency claim.

Limits

One domain, simulated customers. That's why every pilot starts with your own logs.

To run it on your agent's history, Book a call

Join the beta. 10 spots.

If your company runs an AI agent in production, we'd like to compile its most repeated jobs with you.

Book a call

cal.com/agent-compile/beta

To start, we'll ask to see your agent's logs.