An SDK around your agent's model client

The first time, your agent thinks. After that, it's compiled.

Your agent does the same jobs over and over, and thinks each one through from scratch. AgentCompile learns those jobs from your agent's own history and runs them compiled. Everything else goes to your agent, unchanged.

Dataτ-bench retail · 261 tasks, 146 held out

Same answers as your agent. 41% fewer agent calls.

A chevron, your agent, beside an amber block, the compiled job. The block turns slowly and fills a step for each result listed.

  • 41%fewer agent calls, with the same answers.74.5% correct against 73.4%, worst case −0.9 pts at 95%
  • 36%less spent on agent tokens, measured on a Claude agent.Claude Sonnet 4.5 · 21 paired tasks · 26% with prompt caching over 48 pairs
  • 34%of held-out conversations finished with no agent call.86% correct, against the agent's 87% on the same tasks
  • 0times it did anything different.478 recorded agent conversations it never learned from
  • 5,000+benchmark conversations behind these numbers.Gemini 2.5 Flash, Gemini 2.5 Pro and Claude Sonnet 4.5 agents
How we measured

01 / Routing

Where each request goes

AgentCompile wraps your agent's model client. Known jobs run compiled. Anything new, unclear or unusual goes to your model, unchanged.

Your agent calls its model client, which the AgentCompile SDK wraps. A known job runs compiled, and the reply goes back to your agent without calling your model. Anything new, unclear or unusual goes on to your model provider, unchanged. If AgentCompile has any problem, or a step takes too long, the call goes straight to your model.

Known jobs run compiled.

Your agent's model isn't called for them. AgentCompile learns the job, so it handles new customers, new orders and new details. A job goes live only after it gets your past conversations right.

Different customers ask about different orders in their own words. The same learned job answers each of them, compiled, without calling your model.

Your agent is always the fallback.

Anything new, unclear or unusual goes to your agent, unchanged. When AgentCompile hands a conversation back, your agent gets the whole conversation so far.

A customer asks where order 1042 is, and a compiled job answers. The customer then asks for something new, so the conversation goes to your agent, which gets the whole conversation so far.

Nothing irreversible without a yes.

A compiled job asks your customer before it changes anything, and waits for their explicit yes.

A customer asks to cancel order 2291. The compiled job asks them to confirm, and cancels the order only after they reply yes.

It fails open.

If AgentCompile has any problem, or a step takes too long, the call goes straight to your model.

AgentCompile has a problem, so the request goes straight to your model, which replies as usual.

02 / Results

What we measured

Everything here ran on τ-bench retail, the public customer-service agent benchmark, with a simulated customer: the same agent with and without AgentCompile, in the same time window. None of it is customer data.

Dataτ-bench retail · simulated customers
  • 41%fewer agent calls, with the same answers.Gemini 2.5 Pro agent · 261 tasks, 146 held out · ~1,490 conversations per arm · correct 73.4% → 74.5%, worst case −0.9 pts at 95%
  • 42%fewer agent calls on a Claude agent.Claude Sonnet 4.5 agent · 21 paired tasks · 38% with prompt caching over 48 pairs
  • 36%less spent on agent tokens, on the same Claude agent.measured bill · 21 paired tasks, no prompt caching · 26% with caching over 48 pairs
  • 34%of held-out conversations finished end to end with no agent call.86% correct, against the agent's 87% on the same tasks
  • 11.6%of compiled writes miss the right answer, against 15.6% for the agent alone.68 of 584 compiled writes · 464 of 2,974 agent writes · every compiled write is one the customer confirmed
  • 5repeated jobs found in raw agent logs, with no task list.together they cover 85% of the agent's writes
  • 5,000+ benchmark conversations
  • 3 agent models: Gemini 2.5 Flash, Gemini 2.5 Pro, Claude Sonnet 4.5
  • 808 automated tests

On 478 recorded agent conversations it never learned from, it never did anything different.

Agent logs from τ-bench, retail domain. AgentCompile did exactly what the agent did in 430 of them. In the other 48, it didn't act and left the conversation to the agent.

Dataτ-bench replay · retail domain
  • 430 did exactly what the agent did
  • 48 didn't act, left the conversation to the agent
  • 0 did anything different

430 + 48 + 0 = 478 recorded agent conversations it never learned from

  • did exactly what the agent did · 430
  • left to the agent · 48
  • did anything different · 0, no squares
One square per recorded agent conversation, sorted by outcome.

Limits

  • The replay measures agreement with the recorded agent. Correctness is what the live benchmark measures.
  • It covers one domain: τ-bench retail.
  • A pilot measures your own traffic.

How we measured

  1. τ-bench retail, the public customer-service agent benchmark, commit 59a200c6, measured in our benchmark harness.
  2. Held out: 146 train tasks no job was learned from, alongside the 115-task test split.
  3. The same agent model, simulated customer, prompts and time window, with and without AgentCompile.
  4. Agent calls counted; AgentCompile's own calls, under 4 per conversation, reported separately.
  5. One-sided 95% bounds over tasks. Rounds are pooled, because a single round swings about ±4 points.
  6. The fixes between rounds came from reading failures on the test split, which is why every round also ran on the 146 held-out tasks. It holds at parity there.
  7. Latency isn't compared: it tracks the provider's load.
  8. Limits: one domain and simulated customers. A pilot measures your own traffic.

03 / Integration

One line of code.

client = agentcompile.wrap(client)

The best agentic workflow, unlocked.

Install the SDK and wrap the model client your agent already uses. Pass an AgentCompile key, and a conversation id with each conversation. Your model and your prompts stay exactly as they are.

Any model, any framework.

Works with OpenAI- and Anthropic-compatible clients, in any agent framework that lets you pass your own client.

one install · one line · a conversation id per conversation

OpenAI-compatible

from openai import OpenAIAdded line: import agentcompileclient = OpenAI()Added line: client = agentcompile.wrap(Added line:   client, key="<agentcompile key>"Added line: )

Anthropic-compatible

from anthropic import AnthropicAdded line: import agentcompileclient = Anthropic()Added line: client = agentcompile.wrap(Added line:   client, key="<agentcompile key>"Added line: )

also pass, per conversation

Added line: conversation_id="<conversation id>"
Names are placeholders.

Works with

Connected where your agent already is.

The AgentCompile SDK, connected to OpenAI, Anthropic, Google Gemini, Grok, Microsoft Azure and Amazon Bedrock.

04 / Pilot

How a pilot works

Nothing goes live until you've seen how each job did on your history.

  1. 01Learn

    Show us your logs.

    We show you which jobs your agent repeats.

    Your agent's logs: requests like order status, cancelling an order and changing an address come up again and again, and each is marked as a repeated job. A request that doesn't repeat is left unmarked.

  2. 02Prove

    We prove each job on your history.

    Each job has to get your past conversations right.

    A job is replayed on your agent's past conversations. In most it does exactly what your agent did; in the rest it doesn't act and leaves the conversation to your agent. It never does anything different.

  3. 03Run compiled

    Add the SDK.

    Known jobs run compiled. Everything else goes to your agent as usual.

    Live requests at your agent's model client. Known jobs like order status or cancelling an order run compiled, asking the customer for a yes before changing anything. A new request goes to your agent, unchanged.

05 / FAQ

Questions

No. A cache replays old answers. AgentCompile learns the job, so it handles new customers, new orders and new details.

No. Your model stays exactly as it is. AgentCompile decides when a call to it isn't needed.

For anything new, unclear or unusual, the same request it gets today, unchanged. When AgentCompile hands a conversation back, your agent gets the whole conversation so far. For known jobs, your agent's model isn't called.

From your agent's own history: its logs. A job goes live only after it gets your past conversations right, and you see how each job did before it goes live.

To start, your agent's logs. We show you which jobs repeat and prove each one on your history. To go live: add the SDK with an AgentCompile key, and pass a conversation id per conversation.

No. Every result here ran on τ-bench, the public customer-service agent benchmark, retail domain, with a simulated customer. A pilot measures your own traffic.

We go through data handling, key handling and where AgentCompile runs with you on a call, before you share any logs.

Not yet. We're pre-launch and looking for design partners.

We're looking for design partners.

If your company runs an AI agent in production, we'd like to compile its most repeated jobs with you.

founders@tryagentcompile.com

To start, we'll ask to see your agent's logs.