Results

What we measured.

Everything here ran on τ-bench retail, the public customer-service agent benchmark, with a simulated customer: the same agent with and without AgentCompile, in the same time window. None of it is customer data.

Dataτ-bench retail · simulated customers
  • 58%fewer agent calls, with the same answers.Gemini 2.5 Pro agent · 17.2 → 7.3 agent calls per conversation
  • 42%fewer agent calls on a Claude agent.Claude Sonnet 4.5 agent · 21 paired tasks · 38% with prompt caching over 48 pairs
  • 36%less spent on agent tokens, on the same Claude agent.measured bill · 21 paired tasks, no prompt caching · 26% with caching over 48 pairs
  • 34%of held-out conversations finished end to end with no agent call.86% correct, against the agent's 87% on the same tasks
  • 11.6%of compiled writes miss the right answer, against 15.6% for the agent alone.68 of 584 compiled writes · 464 of 2,974 agent writes · every compiled write is one the customer confirmed
  • 5repeated jobs found in raw agent logs, with no task list.together they cover 85% of the agent's writes
  • 5,000+ benchmark conversations
  • 3 agent models: Gemini 2.5 Flash, Gemini 2.5 Pro, Claude Sonnet 4.5
  • 808 automated tests

With and without AgentCompile.

The same agent model on the same τ-bench retail tasks, run both ways.

Dataτ-bench retail · simulated customers
Dot chart, with AgentCompile against without, on τ-bench retail with simulated customers. Agent calls: Gemini 2.5 Pro agent, 58% fewer, 17.2 to 7.3 per conversation; Claude Sonnet 4.5 agent, 42% fewer; OpenAI agent, an estimate of about 40% fewer. Correct answers, Gemini 2.5 Pro agent: 73.4% without, 74.5% with. Agent token spend, Claude Sonnet 4.5 agent: 36% less on the measured bill; OpenAI agent, an estimate of about 35% less.

On 478 recorded agent conversations it never learned from, it never did anything different.

Agent logs from τ-bench, retail domain. AgentCompile did exactly what the agent did in 430 of them. In the other 48, it didn't act and left the conversation to the agent.

Dataτ-bench replay · retail domain
  • 430 did exactly what the agent did
  • 48 didn't act, left the conversation to the agent
  • 0 did anything different

430 + 48 + 0 = 478 recorded agent conversations it never learned from

  • did exactly what the agent did · 430
  • left to the agent · 48
  • did anything different · 0, no squares
One square per recorded agent conversation, sorted by outcome.

Join the beta. 10 spots.

If your company runs an AI agent in production, we'd like to compile its most repeated jobs with you.

Book a call

cal.com/agent-compile/beta

To start, we'll ask to see your agent's logs.