Key result
0 times it did anything different.
430 + 48 + 0 = 478 recorded agent conversations it never learned from
Setup
- Data
- Agent logs from τ-bench, retail domain
- Conversations
- 478 recorded agent conversations it never learned from
- Customers
- Benchmark agent logs, not customer data
The question
Before anything runs compiled, it has to get past conversations right. So: on conversations it has never seen, does it ever do something the agent wouldn't have?
What we found
On 478 recorded agent conversations it never learned from, AgentCompile did exactly what the agent did in 430 of them. In the other 48, it didn't act and left the conversation to the agent. It never did anything different.
When it isn't sure, it doesn't guess. It hands the conversation to your agent.
Why the 48 matter
Leaving a conversation to the agent costs one model call. Doing the wrong thing costs a customer. The 48 are the system choosing the call.
Limits
- One domain: retail customer service.
- Benchmark agent logs, not customer data. A pilot measures your own traffic.