Research

# Same answers, 58% fewer agent calls on a Gemini 2.5 Pro agent

Benchmark / October 1, 2026 / Krishna Bhatnagar

The same agent, on the same τ-bench retail tasks, with and without AgentCompile. Agent calls fell from 17.2 to 7.3 per conversation, and correct answers held.

Key result: 58% fewer agent calls, with the same answers. 17.2 → 7.3 agent calls per conversation

## Setup

- Benchmark: τ-bench retail, the public customer-service agent benchmark
- Customers: Simulated. None of it is customer data.
- Agent model: Gemini 2.5 Pro
- Comparison: The same agent with and without AgentCompile, in the same time window
- Tasks: Held-out test tasks it never learned from

## The question

An agent that does the same jobs again and again thinks each one through from scratch. How many of its model calls does it actually need? We ran the same agent on the same tasks twice: once as it is, and once with AgentCompile around its model client.

## What we found

Agent calls fell from 17.2 to 7.3 per conversation, 58% fewer. Correct answers held: 73.4% without AgentCompile and 74.5% with it. At 95%, the worst case is 0.9 points below the agent alone.

The calls that went away were jobs the agent had already done many times. Anything new, unclear or unusual still went to the agent, unchanged, with the whole conversation so far.

> Fewer calls to the model, and the same answers to the customer.

## Limits

- One domain: retail customer service.
- Simulated customers, not customer data. A pilot measures your own traffic.
- Fewer agent calls is the result. We don't turn it into a cost or a latency claim.
