Everything here ran on τ-bench retail, the public customer-service agent benchmark, with a simulated customer: the same agent with and without AgentCompile, in the same time window. None of it is customer data.
- 58%fewer agent calls, with the same answers.Gemini 2.5 Pro agent · 17.2 → 7.3 agent calls per conversation
- 42%fewer agent calls on a Claude agent.Claude Sonnet 4.5 agent · 21 paired tasks · 38% with prompt caching over 48 pairs
- 36%less spent on agent tokens, on the same Claude agent.measured bill · 21 paired tasks, no prompt caching · 26% with caching over 48 pairs
- 34%of held-out conversations finished end to end with no agent call.86% correct, against the agent's 87% on the same tasks
- 11.6%of compiled writes miss the right answer, against 15.6% for the agent alone.68 of 584 compiled writes · 464 of 2,974 agent writes · every compiled write is one the customer confirmed
- 5repeated jobs found in raw agent logs, with no task list.together they cover 85% of the agent's writes
Replay on recorded conversations
We also replayed 478 recorded agent conversations that AgentCompile never learned from. It did exactly what the agent did in 430, left the other 48 to the agent, and never did anything different.
How to read these numbers
- Every number comes from τ-bench retail, with simulated customers. A pilot measures your own traffic.
- Fewer agent calls is the result. It is not a cost or a latency claim.
- The numbers are quoted as measured, without rounding.
What each number means
- Agent calls: how many times your agent's model is called, per conversation. Fewer means the repeated jobs didn't need it.
- Correct answers: whether the task ended the way the benchmark says it should. This is how we check that fewer calls didn't cost quality.
- Writes: steps that change something. We report how often they miss, for compiled writes and for the agent alone.
- Conversations with no agent call: whole conversations handled end to end by compiled jobs.
How we kept it fair
- Held out: learned only from training tasks, measured on test tasks it never saw.
- The same agent model, simulated customer, prompts and time window, with and without AgentCompile.
- Every model call counted, AgentCompile's own included.
Each result has its own report: the setup, the numbers and the limits. Read the research