All research

Benchmark1 min read

Compiled writes miss less often than the agent alone

A write changes something: a refund, an exchange, a new address. Compiled writes missed the right answer 11.6% of the time, against 15.6% for the agent alone. And 34% of conversations finished with no agent call at all.

Key result

11.6% of compiled writes miss the right answer, against 15.6% for the agent alone.

68 of 584 compiled writes · 464 of 2,974 agent writes

Setup

Benchmark
τ-bench retail, with simulated customers
Writes
584 compiled writes and 2,974 agent writes
Confirmation
Every compiled write is one the customer confirmed

The question

Reading is cheap to get wrong. Writing isn't: a write changes something for the customer. When a job runs compiled, are its writes as good as the agent's?

What we found

68 of 584 compiled writes missed the right answer, 11.6%. For the agent alone, 464 of 2,974 did, 15.6%.

Every compiled write is one the customer confirmed. Nothing irreversible happens without the customer's explicit yes.

Whole conversations

34% of held-out conversations finished end to end with no agent call. Those conversations were 86% correct, against the agent's 87% on the same tasks.

Limits

  • One domain, with simulated customers. A pilot measures your own traffic.
  • A miss is measured against the benchmark's right answer for the task.

Join the beta. 10 spots.

If your company runs an AI agent in production, we'd like to compile its most repeated jobs with you.

Book a call

cal.com/agent-compile/beta

To start, we'll ask to see your agent's logs.