Blog
Notes from AgentCompile.
What we're building, what we measured, and how to add it to your agent.
Introducing AgentCompile
Your agent does the same jobs over and over, and thinks each one through from scratch. AgentCompile learns those jobs from its own history and runs them compiled: same answers, far fewer model calls.
Read the postHow to reduce LLM calls in an AI agent
An agent calls its model on every turn of its loop. Here is what each common fix actually cuts, which ones make calls go away, and where to start.
τ-bench, explained: the benchmark behind our numbers
Every number on this site comes from τ-bench retail. What the benchmark is, how it grades an agent, why we use it, and what it can't tell you.
Prompt caching, model routing, a semantic cache or compiling: what each one cuts
Four ways teams try to make agents cheaper and faster, side by side: what each one does, what it leaves alone, and which ones work together.
Where AgentCompile fits, and where it doesn't on purpose
Support desks, SaaS workflows, phone lines, data teams, back offices. The industry changes; the shape of the work doesn't. A map of where AgentCompile works, where it's heading, and what we leave alone.
Customer support and operations agents
Refunds, exchanges, cancellations, account updates. Customers ask in a thousand ways; the jobs underneath are a short list.
Workflow agents across SaaS tools
Email, calendar, CRM and ticketing. Schedule, update, file, reply: the same shape every time.
Voice agents
Booking, rescheduling, checking an order, changing an account. The same jobs, out loud.
Data and SQL agents
Weekly revenue by region, every Monday, with new dates. The question repeats, so the query can too.
Document and extraction agents
Invoices, receipts and contracts into fields. Reading stays with the model; what comes after is the repeated part.
Internal ops and IT agents
Reset access, provision an account, rotate a key. Lookups, a write, and usually an approval.
Computer-use and desktop agents
The same jobs, clicked through native apps. Where AgentCompile is heading.
Multi-agent handoffs
One agent chooses which specialist takes over. The choice repeats, and so does each specialist's work.
Your agent needs muscle memory
A pianist doesn't think about scales. Your agent still thinks through every refund. On skills that become automatic, and why agents don't have them yet.
What your agent does all day
Agentic work looks open-ended from the outside. Read the logs and most of it is the same handful of jobs, done again and again.
Trust is earned on your own history
Before a job runs without your agent, it has to get your past conversations right. Why we start from history, and what stays protected.
58% fewer agent calls, with the same answers
What we measured on τ-bench retail, how we measured it, and how to read the numbers.
Not a cache. Not fine-tuning.
Two things AgentCompile gets mistaken for, and what it is instead.
Adding AgentCompile in one line
Install the SDK, wrap the client your agent already uses, and give each conversation its id.
Join the beta. 10 spots.
If your company runs an AI agent in production, we'd like to compile its most repeated jobs with you.