Results
What we measured.
Everything here ran on τ-bench retail, the public customer-service agent benchmark, with a simulated customer: the same agent with and without AgentCompile, in the same time window. None of it is customer data.
Dataτ-bench retail · simulated customers
- 58%fewer agent calls, with the same answers.Gemini 2.5 Pro agent · 17.2 → 7.3 agent calls per conversation
- 42%fewer agent calls on a Claude agent.Claude Sonnet 4.5 agent · 21 paired tasks · 38% with prompt caching over 48 pairs
- 36%less spent on agent tokens, on the same Claude agent.measured bill · 21 paired tasks, no prompt caching · 26% with caching over 48 pairs
- 34%of held-out conversations finished end to end with no agent call.86% correct, against the agent's 87% on the same tasks
- 11.6%of compiled writes miss the right answer, against 15.6% for the agent alone.68 of 584 compiled writes · 464 of 2,974 agent writes · every compiled write is one the customer confirmed
- 5repeated jobs found in raw agent logs, with no task list.together they cover 85% of the agent's writes
- 5,000+ benchmark conversations
- 3 agent models: Gemini 2.5 Flash, Gemini 2.5 Pro, Claude Sonnet 4.5
- 808 automated tests
With and without AgentCompile.
The same agent model on the same τ-bench retail tasks, run both ways.
Dataτ-bench retail · simulated customers
On 478 recorded agent conversations it never learned from, it never did anything different.
Agent logs from τ-bench, retail domain. AgentCompile did exactly what the agent did in 430 of them. In the other 48, it didn't act and left the conversation to the agent.
Dataτ-bench replay · retail domain
- 430 did exactly what the agent did
- 48 didn't act, left the conversation to the agent
- 0 did anything different
430 + 48 + 0 = 478 recorded agent conversations it never learned from
- did exactly what the agent did · 430
- left to the agent · 48
- did anything different · 0, no squares
Join the beta. 10 spots.
If your company runs an AI agent in production, we'd like to compile its most repeated jobs with you.