Research

# How we measure

Method / October 1, 2026 / Krishna Bhatnagar

Held out, same agent both ways, every model call counted, AgentCompile's own included. The rules every number on this site follows.

Key result: 5,000+ benchmark conversations. 3 agent models: Gemini 2.5 Flash, Gemini 2.5 Pro, Claude Sonnet 4.5 · 808 automated tests

## Setup

- Benchmark: τ-bench, the public customer-service agent benchmark, retail domain
- Held out: Learned only from training tasks, measured on test tasks it never saw
- Same both ways: Agent model, simulated customer, prompts and time window
- Counted: Every model call, AgentCompile's own included
- Reported: Sample sizes and 95% intervals

## The rules

- Held out. AgentCompile learns only from training tasks and is measured on test tasks it never saw.
- Same agent both ways. The agent model, the simulated customer, the prompts and the time window are the same with and without AgentCompile.
- Every model call counted, AgentCompile's own included.
- Sample sizes and 95% intervals, with every number.
- Per-task results are available. Ask us.

## How to read our numbers

- Every number keeps its source: τ-bench retail, with simulated customers.
- None of it is customer data. A pilot measures your own traffic.
- Fewer agent calls is a result on its own. Don't turn it into a cost or a latency claim.

## Limits

One domain, simulated customers. That's why every pilot starts with your own logs.

To run it on your agent's history, [Book a call](https://cal.com/agent-compile/beta)
