Key result
42% fewer agent calls on a Claude agent.
Claude Sonnet 4.5 agent · 21 paired tasks
Setup
- Benchmark
- τ-bench retail, with simulated customers
- Agent model
- Claude Sonnet 4.5
- Pairs
- 21 paired tasks without prompt caching; 48 pairs with it
- Spend
- The measured bill for agent tokens
The question
Does it hold on a different model, and does the bill follow the calls? We ran a Claude Sonnet 4.5 agent both ways, task by task, and read the bill for its tokens.
What we found
On 21 paired tasks without prompt caching, agent calls fell 42% and the measured bill for agent tokens fell 36%.
With prompt caching on, over 48 pairs, agent calls fell 38% and spend fell 26%. Both ways, the agent made fewer calls and spent less.
Limits
- Small samples: 21 pairs, and 48 with prompt caching.
- One domain, with simulated customers. A pilot measures your own traffic.