Key result
11.6% of compiled writes miss the right answer, against 15.6% for the agent alone.
68 of 584 compiled writes · 464 of 2,974 agent writes
Setup
- Benchmark
- τ-bench retail, with simulated customers
- Writes
- 584 compiled writes and 2,974 agent writes
- Confirmation
- Every compiled write is one the customer confirmed
The question
Reading is cheap to get wrong. Writing isn't: a write changes something for the customer. When a job runs compiled, are its writes as good as the agent's?
What we found
68 of 584 compiled writes missed the right answer, 11.6%. For the agent alone, 464 of 2,974 did, 15.6%.
Every compiled write is one the customer confirmed. Nothing irreversible happens without the customer's explicit yes.
Whole conversations
34% of held-out conversations finished end to end with no agent call. Those conversations were 86% correct, against the agent's 87% on the same tasks.
Limits
- One domain, with simulated customers. A pilot measures your own traffic.
- A miss is measured against the benchmark's right answer for the task.