Hi — api.py's client.messages.create(...) sends tools and the whole message list with no cache_control, so each agent turn re-buys the entire prefix (tool schemas + conversation so far) at full input rate. Marking the last tool and the latest message content block with {"cache_control": {"type": "ephemeral"}} lets subsequent turns read that shared prefix at ~10% of base input price — in an agentic coding loop that prefix is most of the input bill. Happy to open a PR if useful.
(Context: I build a cost-analysis tool for agent workloads — https://lens.r-lattice.com — and this repo came up when scanning public Claude agents for uncached call sites. No affiliation needed to take the fix.)
Hi — api.py's
client.messages.create(...)sendstoolsand the whole message list with nocache_control, so each agent turn re-buys the entire prefix (tool schemas + conversation so far) at full input rate. Marking the last tool and the latest message content block with{"cache_control": {"type": "ephemeral"}}lets subsequent turns read that shared prefix at ~10% of base input price — in an agentic coding loop that prefix is most of the input bill. Happy to open a PR if useful.(Context: I build a cost-analysis tool for agent workloads — https://lens.r-lattice.com — and this repo came up when scanning public Claude agents for uncached call sites. No affiliation needed to take the fix.)