Research

# Compiled writes miss less often than the agent alone

Benchmark / October 1, 2026 / Krishna Bhatnagar

A write changes something: a refund, an exchange, a new address. Compiled writes missed the right answer 11.6% of the time, against 15.6% for the agent alone. And 34% of conversations finished with no agent call at all.

Key result: 11.6% of compiled writes miss the right answer, against 15.6% for the agent alone. 68 of 584 compiled writes · 464 of 2,974 agent writes

## Setup

- Benchmark: τ-bench retail, with simulated customers
- Writes: 584 compiled writes and 2,974 agent writes
- Confirmation: Every compiled write is one the customer confirmed

## The question

Reading is cheap to get wrong. Writing isn't: a write changes something for the customer. When a job runs compiled, are its writes as good as the agent's?

## What we found

68 of 584 compiled writes missed the right answer, 11.6%. For the agent alone, 464 of 2,974 did, 15.6%.

Every compiled write is one the customer confirmed. Nothing irreversible happens without the customer's explicit yes.

## Whole conversations

34% of held-out conversations finished end to end with no agent call. Those conversations were 86% correct, against the agent's 87% on the same tasks.

## Limits

- One domain, with simulated customers. A pilot measures your own traffic.
- A miss is measured against the benchmark's right answer for the task.
