All posts

Guides2 min read

How to reduce LLM calls in an AI agent

An agent calls its model on every turn of its loop. Here is what each common fix actually cuts, which ones make calls go away, and where to start.

An AI agent doesn't call its model once per conversation. It calls it on every turn of its loop: read the message, decide on a tool, read the result, decide again, reply. Each of those calls reads the whole conversation so far. That is why agents cost more and take longer than the chat demo that inspired them.

There are several ways to bring that down. They are often lumped together, but they cut different things. Some make each call cheaper. Only a few make calls go away.

Where the calls come from

  • Every step of the loop: each tool decision and each reply is a model call.
  • Reading tool results: after a lookup, the model is called again to read what came back.
  • Retries: a malformed tool call or a timeout means the same step runs again.
  • Repetition: the same job, done for a different customer, is reasoned out from scratch every time.

Five levers, and what each one cuts

  1. Trim the context. Send each call less: shorter instructions, fewer old turns. Each call gets cheaper; the number of calls stays the same.
  2. Turn on prompt caching. Providers can reuse a repeated prefix, like your system prompt, across calls. Each call gets cheaper; the number of calls stays the same.
  3. Route to a smaller model. Send easy steps to a cheaper model. Each call gets cheaper; the number of calls stays the same.
  4. Add a semantic cache. Replay an earlier answer when a similar question comes back. Calls go away, but an agent that changes things can't safely replay an old action for a new order.
  5. Compile the repeated jobs. Jobs your agent does again and again run without calling its model, with the new details filled in. Calls go away, and anything new still goes to your agent.

Most levers make each call cheaper. Only a few make calls go away.

What compiling did on a benchmark

On τ-bench retail, the public customer-service agent benchmark, with simulated customers, compiling the repeated jobs meant 58% fewer agent calls: 17.2 to 7.3 per conversation, with the same answers. None of it is customer data. A pilot measures your own traffic.

Where to start

  1. Count model calls per conversation, not per day. That is the number every lever moves.
  2. Read your logs for jobs that repeat. Most production agents spend most of their calls on a short list of them.
  3. Trim context and turn on prompt caching. Both are cheap to do and keep working alongside everything else.
  4. Take the repeated jobs off the model's plate, and keep your agent as the fallback for everything else.

Questions we get

Does prompt caching reduce the number of calls?
No. It makes the repeated part of each call cheaper. Your agent still calls the model on every turn.
Is model routing the same as compiling?
No. Routing still calls a model, just a smaller one. A compiled job doesn't call your agent's model at all.
Can these be combined?
Yes. AgentCompile wraps the client your agent already uses, so prompt caching and your choice of model keep working for every call that still goes to your model.

Add AgentCompile to your agent in one line. Open the Cookbook

Join the beta. 10 spots.

If your company runs an AI agent in production, we'd like to compile its most repeated jobs with you.

Book a call

cal.com/agent-compile/beta

To start, we'll ask to see your agent's logs.