Key result
58% fewer agent calls, with the same answers.
17.2 → 7.3 agent calls per conversation
Setup
- Benchmark
- τ-bench retail, the public customer-service agent benchmark
- Customers
- Simulated. None of it is customer data.
- Agent model
- Gemini 2.5 Pro
- Comparison
- The same agent with and without AgentCompile, in the same time window
- Tasks
- Held-out test tasks it never learned from
The question
An agent that does the same jobs again and again thinks each one through from scratch. How many of its model calls does it actually need? We ran the same agent on the same tasks twice: once as it is, and once with AgentCompile around its model client.
What we found
Agent calls fell from 17.2 to 7.3 per conversation, 58% fewer. Correct answers held: 73.4% without AgentCompile and 74.5% with it. At 95%, the worst case is 0.9 points below the agent alone.
The calls that went away were jobs the agent had already done many times. Anything new, unclear or unusual still went to the agent, unchanged, with the whole conversation so far.
Fewer calls to the model, and the same answers to the customer.
Limits
- One domain: retail customer service.
- Simulated customers, not customer data. A pilot measures your own traffic.
- Fewer agent calls is the result. We don't turn it into a cost or a latency claim.