All research

Replay1 min read

478 recorded conversations it never learned from, and it never did anything different

On agent logs it had never seen, AgentCompile did exactly what the agent did in 430 conversations, left 48 to the agent, and did something different in none.

Key result

0 times it did anything different.

430 + 48 + 0 = 478 recorded agent conversations it never learned from

Setup

Data
Agent logs from τ-bench, retail domain
Conversations
478 recorded agent conversations it never learned from
Customers
Benchmark agent logs, not customer data

The question

Before anything runs compiled, it has to get past conversations right. So: on conversations it has never seen, does it ever do something the agent wouldn't have?

What we found

On 478 recorded agent conversations it never learned from, AgentCompile did exactly what the agent did in 430 of them. In the other 48, it didn't act and left the conversation to the agent. It never did anything different.

When it isn't sure, it doesn't guess. It hands the conversation to your agent.

Why the 48 matter

Leaving a conversation to the agent costs one model call. Doing the wrong thing costs a customer. The 48 are the system choosing the call.

Limits

  • One domain: retail customer service.
  • Benchmark agent logs, not customer data. A pilot measures your own traffic.

Join the beta. 10 spots.

If your company runs an AI agent in production, we'd like to compile its most repeated jobs with you.

Book a call

cal.com/agent-compile/beta

To start, we'll ask to see your agent's logs.