Research
What we measured, in full.
One report per result: the question, the setup, the numbers and the limits. Every number keeps its source.
Same answers, 58% fewer agent calls on a Gemini 2.5 Pro agent
The same agent, on the same τ-bench retail tasks, with and without AgentCompile. Agent calls fell from 17.2 to 7.3 per conversation, and correct answers held.
Read the reportA Claude agent: 42% fewer agent calls, 36% less on the bill
The same test on a Claude Sonnet 4.5 agent, with the bill measured. Fewer calls and less spent on agent tokens, with prompt caching and without it.
Read the report478 recorded conversations it never learned from, and it never did anything different
On agent logs it had never seen, AgentCompile did exactly what the agent did in 430 conversations, left 48 to the agent, and did something different in none.
Read the reportCompiled writes miss less often than the agent alone
A write changes something: a refund, an exchange, a new address. Compiled writes missed the right answer 11.6% of the time, against 15.6% for the agent alone. And 34% of conversations finished with no agent call at all.
Read the reportFive repeated jobs cover 85% of an agent's writes
Given raw agent logs and no task list, AgentCompile found 5 repeated jobs. Together they cover 85% of what the agent changes.
Read the reportHow we measure
Held out, same agent both ways, every model call counted, AgentCompile's own included. The rules every number on this site follows.
Read the report
Join the beta. 10 spots.
If your company runs an AI agent in production, we'd like to compile its most repeated jobs with you.