Tuesday Sep 8

Anthropic Cuts Cache Reads By 75%

8SEP
$1.00$0.25CACHE READSAGENTS WIN

Reading a cached token now costs 25 cents per million instead of a dollar. Input and output prices did not move. The whole cut lands on the thing agents do most.

A cache read is the model re-reading context it already saw. Repo, system prompt, tool specs, prior turns. Long agent runs do it constantly.

A typical workload gets about 25% cheaper. A context-heavy agentic one drops closer to 45%. Input stays $10 per million, output $50.

It shipped with Claude Fable 5.1 and Mythos 5.1 on September 1. Cache writes are unchanged at $12.50 per million.

full brief & sources

⚡ Why this matters

  • Pricing moved on one line item, and it is the line item that decides whether long-running agents are affordable.
  • Cutting cache reads and nothing else is a bet that context, not generation, is where the spend went.
  • If your agent cost model is built on input and output rates, it is now wrong.

🔍 What happened

  • Anthropic released Claude Fable 5.1 and Mythos 5.1 on September 1.
  • Cache reads dropped from $1.00 to $0.25 per million tokens, a 75% cut.
  • Standard rates are unchanged: $10 per million input, $50 per million output.
  • Cache writes stay at $12.50 per million for the five-minute cache.
  • A typical workload comes out roughly 25% cheaper overall.
  • A context-heavy agentic workload, where cache reads dominate spend, drops closer to 45%.

💬 Smart takes

  • Anthropic's framing: the cut targets persistent work, where the same repo, instructions and tool specs get resent turn after turn.
  • Enterprise DNA: the 'cheaper' claim does not hold at real task-level cost once you account for how the model is actually used.
  • Skeptic: a price cut on the fastest-growing usage line is a volume play, not generosity. Total bills can still go up.

🧭 Where this goes

  1. Likelyrivals match the cache-read price within a quarter, because it is now the comparison buyers run.
  2. Likelyagent frameworks start optimising for cache-hit rate the way they once optimised for prompt length.
  3. Possible'cost per pull request' replaces 'cost per million tokens' as the number engineering leads quote.
  4. Possiblesomeone publishes a benchmark showing the 45% claim only holds on a narrow workload shape.
  5. Wild Cardcache reads go effectively free and pricing shifts entirely to output, which changes how agents get designed.

🥄 The Spoon Take

Model launches used to be about the benchmark. This one is about the invoice. The interesting move is which line they cut: not the clever tokens, the boring re-read ones. That tells you where the money was actually going.

🤔 Pushback

Independent analysis says the cheaper claim does not survive contact with real task-level cost, and a lower unit price on a growing workload can still mean a bigger bill.