An SDK around your agent's model client

# The first time, your agent thinks. After that, it's compiled.

Your agent does the same jobs over and over, and thinks each one through from scratch. AgentCompile learns those jobs from your agent's own history and runs them compiled. Everything else goes to your agent, unchanged.

[data] τ-bench retail · 261 tasks, 146 held out

## 41% fewer agent calls.

Same answers as your agent.

A chevron, your agent, beside an amber block, the compiled job. The block turns slowly and fills a step for each result listed.

- 41% fewer agent calls, with the same answers. 74.5% correct against 73.4%, worst case −0.9 pts at 95%
- 36% less spent on agent tokens, measured on a Claude agent. Claude Sonnet 4.5 · 21 paired tasks · 26% with prompt caching over 48 pairs
- 34% of held-out conversations finished with no agent call. 86% correct, against the agent's 87% on the same tasks
- 0 times it did anything different. 478 recorded agent conversations it never learned from
- 5,000+ benchmark conversations behind these numbers. Gemini 2.5 Flash, Gemini 2.5 Pro and Claude Sonnet 4.5 agents

01 / Routing

## Where each request goes.

AgentCompile wraps your agent's model client. Known jobs run compiled. Anything new, unclear or unusual goes to your model, unchanged.

Your agent calls its model client, which the AgentCompile SDK wraps. A known job runs compiled, and the reply goes back to your agent without calling your model. Anything new, unclear or unusual goes on to your model provider, unchanged. If AgentCompile has any problem, or a step takes too long, the call goes straight to your model.

Four rules for every request. Known jobs run compiled, your agent is always the fallback, nothing irreversible happens without a yes, and it fails open.

### Known jobs run compiled.

Your agent's model isn't called for them. AgentCompile learns the job, so it handles new customers, new orders and new details. A job goes live only after it gets your past conversations right.

Three different requests run into one compiled job, and three identical answers come out. Your model stands beside it and is never called.

### Your agent is always the fallback.

Anything new, unclear or unusual goes to your agent, unchanged. When AgentCompile hands a conversation back, your agent gets the whole conversation so far.

A known request runs into the compiled job. A new one, carrying the conversation so far, turns off before the job and goes to your agent.

### Nothing irreversible without a yes.

A compiled job asks your customer before it changes anything, and waits for their explicit yes.

A change from the compiled job waits at a lowered bar. A yes arrives, the bar lifts, and only then does the change go through.

### It fails open.

If AgentCompile has any problem, or a step takes too long, the call goes straight to your model.

The compiled job has a problem, so the request takes the path around it, straight to your model.

02 / Results

## What we measured.

Everything here ran on τ-bench retail, the public customer-service agent benchmark, with a simulated customer: the same agent with and without AgentCompile, in the same time window. None of it is customer data.

[data] τ-bench retail · simulated customers

- [data] 41% fewer agent calls, with the same answers. Gemini 2.5 Pro agent · 261 tasks, 146 held out · ~1,490 conversations per arm · correct 73.4% → 74.5%, worst case −0.9 pts at 95%
- [data] 42% fewer agent calls on a Claude agent. Claude Sonnet 4.5 agent · 21 paired tasks · 38% with prompt caching over 48 pairs
- [data] 36% less spent on agent tokens, on the same Claude agent. measured bill · 21 paired tasks, no prompt caching · 26% with caching over 48 pairs
- [data] 34% of held-out conversations finished end to end with no agent call. 86% correct, against the agent's 87% on the same tasks
- [data] 11.6% of compiled writes miss the right answer, against 15.6% for the agent alone. 68 of 584 compiled writes · 464 of 2,974 agent writes · every compiled write is one the customer confirmed
- [data] 5 repeated jobs found in raw agent logs, with no task list. together they cover 85% of the agent's writes

- 5,000+ benchmark conversations
- 3 agent models: Gemini 2.5 Flash, Gemini 2.5 Pro, Claude Sonnet 4.5
- 808 automated tests

### On 478 recorded agent conversations it never learned from, it never did anything different.

Agent logs from τ-bench, retail domain. AgentCompile did exactly what the agent did in 430 of them. In the other 48, it didn't act and left the conversation to the agent.

[data] τ-bench replay · retail domain

- 430 did exactly what the agent did
- 48 didn't act, left the conversation to the agent
- 0 did anything different

430 + 48 + 0 = 478 recorded agent conversations it never learned from

Chart of 478 recorded agent conversations, one square each, sorted by outcome: 430 amber squares where AgentCompile did exactly what the agent did, 48 outlined squares where it didn't act and left the conversation to the agent, and no squares for did anything different, because that count is 0.

- did exactly what the agent did · 430
- left to the agent · 48
- did anything different · 0, no squares

One square per recorded agent conversation, sorted by outcome.

03 / Integration

## One line of code.

`client = agentcompile.wrap(client)`

The best agentic workflow, unlocked.

Install the SDK and wrap the model client your agent already uses. Pass an AgentCompile key, and a conversation id with each conversation. Your model and your prompts stay exactly as they are.

### Any model, any framework.

Works with OpenAI- and Anthropic-compatible clients, in any agent framework that lets you pass your own client.

one install · one line · a conversation id per conversation

OpenAI-compatible

```diff
  from openai import OpenAI
+ import agentcompile

  client = OpenAI()
+ client = agentcompile.wrap(
+   client, key="<agentcompile key>"
+ )
```

Anthropic-compatible

```diff
  from anthropic import Anthropic
+ import agentcompile

  client = Anthropic()
+ client = agentcompile.wrap(
+   client, key="<agentcompile key>"
+ )
```

also pass, per conversation

```diff
+ conversation_id="<conversation id>"
```

Names are placeholders.

04 / Works with

## Connected where your agent already is.

AgentCompile wraps the model client your agent already uses, so it goes wherever that client goes: any model behind an OpenAI- or Anthropic-compatible client. Your model and your prompts stay exactly as they are.

The AgentCompile SDK, connected to OpenAI, Anthropic, Google Gemini, Grok, Microsoft Azure and Amazon Bedrock.

- OpenAI
- Anthropic
- Google Gemini
- Grok
- Microsoft Azure
- Amazon Bedrock

05 / Pilot

## How a pilot works.

Nothing goes live until you've seen how each job did on your history.

### 01 Learn: Show us your logs.

We show you which jobs your agent repeats.

Your agent's logs: requests like order status, cancelling an order and changing an address come up again and again, and each is marked as a repeated job. A request that doesn't repeat is left unmarked.

### 02 Prove: We prove each job on your history.

Each job has to get your past conversations right.

A job is replayed on your agent's past conversations. In most it does exactly what your agent did; in the rest it doesn't act and leaves the conversation to your agent. It never does anything different.

### 03 Run compiled: Add the SDK.

Known jobs run compiled. Everything else goes to your agent as usual.

One line adds the SDK, wrapping your agent's model client. Then live requests at that client: known jobs like order status or cancelling an order run compiled, asking the customer for a yes before changing anything. A new request goes to your agent, unchanged.

06 / FAQ

## Questions.

What it is, why use it, and what it needs from you.

### What is AgentCompile?

An SDK that wraps your agent's model client. It learns the jobs your agent repeats from your agent's own history, proves each one on that history, and then runs them compiled, without calling your agent's model. Anything new, unclear or unusual goes to your agent, unchanged.

### Why use it?

Your agent does the same jobs over and over, and thinks each one through from scratch. Once a job is compiled, your agent's model isn't called for it. On τ-bench retail, that meant 41% fewer agent calls, with the same answers.

### Who is it for?

Companies running an AI agent in production, and the engineers who run it.

### How do I add it?

Install the SDK and wrap the model client your agent already uses. Pass an AgentCompile key, and a conversation id with each conversation. Your model and your prompts stay exactly as they are.

### Which models does it work with?

Any model behind an OpenAI- or Anthropic-compatible client, including OpenAI, Anthropic, Google Gemini, Grok, Microsoft Azure and Amazon Bedrock, in any agent framework that lets you pass your own client.

### What if AgentCompile goes down?

It fails open. If AgentCompile has any problem, or a step takes too long, the call goes straight to your model.

### Is this a cache?

No. A cache replays old answers. AgentCompile learns the job, so it handles new customers, new orders and new details.

### Is this fine-tuning?

No. Your model stays exactly as it is. AgentCompile decides when a call to it isn't needed.

### What does my agent see?

For anything new, unclear or unusual, the same request it gets today, unchanged. When AgentCompile hands a conversation back, your agent gets the whole conversation so far. For known jobs, your agent's model isn't called.

### Where do the jobs come from?

From your agent's own history: its logs. A job goes live only after it gets your past conversations right, and you see how each job did before it goes live.

### What do you need from us?

To start, your agent's logs. We show you which jobs repeat and prove each one on your history. To go live: add the SDK with an AgentCompile key, and pass a conversation id per conversation.

### Was the evidence measured on customer data?

No. Every result here ran on τ-bench, the public customer-service agent benchmark, retail domain, with a simulated customer. A pilot measures your own traffic.

### What about our logs, our provider key and where it runs?

We go through data handling, key handling and where AgentCompile runs with you on a call, before you share any logs.

### Can I use it today?

Not yet. We're pre-launch and looking for design partners.

## There are compilers for code. There are compilers for models. This is the compiler for agents.

## We're looking for design partners.

If your company runs an AI agent in production, we'd like to compile its most repeated jobs with you.

`founders@tryagentcompile.com`

To start, we'll ask to see your agent's logs.

---

The compiler for agents.

- [data] measured.

© 2026 AgentCompile
