Key result
5,000+ benchmark conversations.
3 agent models: Gemini 2.5 Flash, Gemini 2.5 Pro, Claude Sonnet 4.5 · 808 automated tests
Setup
- Benchmark
- τ-bench, the public customer-service agent benchmark, retail domain
- Held out
- Learned only from training tasks, measured on test tasks it never saw
- Same both ways
- Agent model, simulated customer, prompts and time window
- Counted
- Every model call, AgentCompile's own included
- Reported
- Sample sizes and 95% intervals
The rules
- Held out. AgentCompile learns only from training tasks and is measured on test tasks it never saw.
- Same agent both ways. The agent model, the simulated customer, the prompts and the time window are the same with and without AgentCompile.
- Every model call counted, AgentCompile's own included.
- Sample sizes and 95% intervals, with every number.
- Per-task results are available. Ask us.
How to read our numbers
- Every number keeps its source: τ-bench retail, with simulated customers.
- None of it is customer data. A pilot measures your own traffic.
- Fewer agent calls is a result on its own. Don't turn it into a cost or a latency claim.
Limits
One domain, simulated customers. That's why every pilot starts with your own logs.
To run it on your agent's history, Book a call