# What we measured, in full.

One report per result: the question, the setup, the numbers and the limits. Every number keeps its source.

- [Same answers, 58% fewer agent calls on a Gemini 2.5 Pro agent](/research/58-percent-fewer-agent-calls-gemini.md): October 1, 2026. The same agent, on the same τ-bench retail tasks, with and without AgentCompile. Agent calls fell from 17.2 to 7.3 per conversation, and correct answers held.
- [A Claude agent: 42% fewer agent calls, 36% less on the bill](/research/claude-sonnet-4-5-agent.md): October 1, 2026. The same test on a Claude Sonnet 4.5 agent, with the bill measured. Fewer calls and less spent on agent tokens, with prompt caching and without it.
- [478 recorded conversations it never learned from, and it never did anything different](/research/478-recorded-conversations.md): October 1, 2026. On agent logs it had never seen, AgentCompile did exactly what the agent did in 430 conversations, left 48 to the agent, and did something different in none.
- [Compiled writes miss less often than the agent alone](/research/compiled-writes.md): October 1, 2026. A write changes something: a refund, an exchange, a new address. Compiled writes missed the right answer 11.6% of the time, against 15.6% for the agent alone. And 34% of conversations finished with no agent call at all.
- [Five repeated jobs cover 85% of an agent's writes](/research/five-repeated-jobs.md): October 1, 2026. Given raw agent logs and no task list, AgentCompile found 5 repeated jobs. Together they cover 85% of what the agent changes.
- [How we measure](/research/how-we-measure.md): October 1, 2026. Held out, same agent both ways, every model call counted, AgentCompile's own included. The rules every number on this site follows.
