All research

Benchmark1 min read

Same answers, 58% fewer agent calls on a Gemini 2.5 Pro agent

The same agent, on the same τ-bench retail tasks, with and without AgentCompile. Agent calls fell from 17.2 to 7.3 per conversation, and correct answers held.

Key result

58% fewer agent calls, with the same answers.

17.2 → 7.3 agent calls per conversation

Setup

Benchmark
τ-bench retail, the public customer-service agent benchmark
Customers
Simulated. None of it is customer data.
Agent model
Gemini 2.5 Pro
Comparison
The same agent with and without AgentCompile, in the same time window
Tasks
Held-out test tasks it never learned from

The question

An agent that does the same jobs again and again thinks each one through from scratch. How many of its model calls does it actually need? We ran the same agent on the same tasks twice: once as it is, and once with AgentCompile around its model client.

What we found

Agent calls fell from 17.2 to 7.3 per conversation, 58% fewer. Correct answers held: 73.4% without AgentCompile and 74.5% with it. At 95%, the worst case is 0.9 points below the agent alone.

The calls that went away were jobs the agent had already done many times. Anything new, unclear or unusual still went to the agent, unchanged, with the whole conversation so far.

Fewer calls to the model, and the same answers to the customer.

Limits

  • One domain: retail customer service.
  • Simulated customers, not customer data. A pilot measures your own traffic.
  • Fewer agent calls is the result. We don't turn it into a cost or a latency claim.

Join the beta. 10 spots.

If your company runs an AI agent in production, we'd like to compile its most repeated jobs with you.

Book a call

cal.com/agent-compile/beta

To start, we'll ask to see your agent's logs.