Skip to main content
TokenCost logoTokenCost
Model ReleaseSeptember 2, 2026ยท12 min read

Claude Fable 5.1 charges what Fable 5 charged on every line of the rate card except one. That one is the cache read, and it is the first time Anthropic has priced a cache hit at anything other than a tenth of input.

Anthropic shipped Fable 5.1 and Mythos 5.1 on September 1. Input is still $10.00 per million tokens, output is still $50.00, both cache writes are still $12.50 and $20.00, and the Batch API still halves the base rates to $5.00 and $25.00. Every headline number is the number it was. The cache read went from $1.00 to $0.25, and because that is the only thing that moved, it is also the only place your bill can fall. How far it falls is a division you can do in your head.

A wall grid of identical dark metal post office box doors with one blank door standing out among them

Photo by MountainAsh on Unsplash

The change, and what follows from it

  • Fable 5.1 and Mythos 5.1 shipped September 1 on Fable 5's exact card except the cache read, which went $1.00 to $0.25.
  • That is a 0.025x multiplier. Anthropic has charged 0.1x for a cache hit on all seventeen rows of its price list until now, and says so in a footnote whose word is standard.
  • Your saving is 0.75 times the share of your Fable 5 bill that was cache reads. Nothing else moved, so nothing else can fall. The ceiling is 75%.
  • Anthropic's 25% and 45% are that formula at cache-read shares of 33.3% and 60%. On a standard 7:2:1 benchmark blend it is 6.82%.
  • Caching still starts paying at the same point, one read on the five-minute cache and two on the one-hour, because writes did not change.
  • Writes did not change, so a cache miss went from costing 12.5 reads to 50, and the one-hour cache went from 1.86x to 3.56x cheaper than the five-minute one over a sustained hour.

Six lines, one of them different

LineFable 5Fable 5.1Change
Base input$10.00$10.00unchanged
5-minute cache write$12.50$12.50unchanged
1-hour cache write$20.00$20.00unchanged
Cache read$1.00$0.25down 75%
Output$50.00$50.00unchanged
Batch input / output$5.00 / $25.00$5.00 / $25.00unchanged

Per million tokens, read from Anthropic's pricing documentation on September 2, 2026. Mythos 5.1 carries the identical card. The 1,000,000-token context window, the 128,000-token output maximum and the 512-token cache minimum are all the same as well, and the full window is still billed flat, with no long-context tier of the kind xAI puts on Grok at 200,000 tokens.

A point-one release that leaves input and output alone is unusual enough on its own. What makes this one worth an article is which line Anthropic chose instead, and the fact that the line has never been touched before on any model the company sells.

A cache hit has been a tenth of input for the entire history of this price list

Anthropic's caching prices have always been multipliers rather than independent figures. A five-minute write is 1.25x base input, a one-hour write is 2x, and a cache read is 0.1x. Those three multipliers hold on every row of the price list, from Haiku 3.5 at $0.80 input and $0.08 a read up through the retired Opus 4.1 at $15.00 and $1.50. Fifteen models before September 1, three constants, no exceptions.

There are two exceptions now, and Anthropic has written them into a footnote under the pricing table: cache hits and refreshes on Fable 5.1 and Mythos 5.1 are priced at 0.025x the base input price, and all other models use the standard 0.1x multiplier. The word standard is doing a lot of work in that sentence. It is the only place on the page where Anthropic acknowledges that the formula it has used for every model it has ever priced now has a hole in it.

The correction has not propagated. Anthropic's models overview, the page that puts Fable 5.1 in its leftmost column as the model to reach for on long-horizon agentic work, carries a footnote under the pricing row reading that Batch API requests are 50% off and prompt cache reads cost 10% of the base input price. For the model that page is recommending, that is four times the truth. The pricing page itself does the same thing more carefully, stating the 10% rule and its one-read break-even in prose and then adding the 2.5% case as the following sentence. A cut this size leaves a lot of documents to update and Anthropic has not finished.

ModelBase inputCache readMultiplier
Claude Haiku 4.5$1.00$0.100.1x
Claude Sonnet 5$2.00$0.200.1x
Claude Fable 5.1$10.00$0.250.025x
Claude Mythos 5.1$10.00$0.250.025x
Claude Sonnet 4.6$3.00$0.300.1x
Claude Opus 5$5.00$0.500.1x
Claude Opus 4.8$5.00$0.500.1x
Claude Fable 5$10.00$1.000.1x
Claude Mythos 5$10.00$1.000.1x
Claude Opus 4.1, retired$15.00$1.500.1x

Ten of the seventeen rows, sorted by what a cache hit costs; the seven left out are Opus 4.7, 4.6, 4.5 and 4, Sonnet 4.5 and 4, and Haiku 3.5, none of which does anything the rows above do not. Read down the middle column and the ordering is the one the left column implies everywhere except positions three and four, where the most expensive model Anthropic sells sits between the $2.00 model and the $3.00 one. Fable 5.1 reads cached tokens for half of what Opus 5 does, on a base price twice as high.

Your discount is your cache-read share, times three quarters

Anthropic quotes two numbers in the announcement: Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, and for highly agentic work the savings will often be much larger, up to approximately 45%. Both figures are usually reported as though they were measurements of something. They are not. Because exactly one line of the card changed, the saving on any workload is forced, and it is a single multiplication.

Cache reads got 75% cheaper and nothing else moved, so your bill falls by 75% of whatever fraction of it was cache reads. Take that fraction, multiply by 0.75, and you have your number. Run it backwards and Anthropic's two figures decode: 25% is the saving for a workload where cache reads were 33.3% of the Fable 5 bill, and 45% is the saving where they were 60% of it. The two headline percentages are not two different claims. They are one claim about two different cache-read shares.

Cache reads as a share of your Fable 5 billWhat you save on Fable 5.1
10%7.50%
25%18.75%
33.3%25.00%
50%37.50%
60%45.00%
80%60.00%
100%75.00%

The bottom row is the ceiling and it is worth naming, because a lot of the coverage this week has described the change as models getting cheaper without a bound on how much cheaper. A workload made entirely of cache reads saves 75% and there is no arrangement of traffic that saves more. If you are paying for output at $50.00 per million, which almost everybody is, you are nowhere near that line.

It is worth seeing how far from 25% an ordinary blend lands. Artificial Analysis prices models on a fixed 7:2:1 mix of cache reads, fresh input and output. Run Fable 5 through it and you get $0.70 plus $2.00 plus $5.00, or $7.70 per million. Run Fable 5.1 and the first term becomes $0.175, for $7.175. That is a saving of 6.82%, because on that mix cache reads are only 9.09% of the bill, and 9.09% times 0.75 is 6.82%. Seven out of every ten tokens in that blend are cache reads and the discount is still under seven percent, because the output tokens nobody caches are five times the price of the input ones.

So you do not need Anthropic's estimate at all. Your last invoice has a cache-read line on it. Divide it by the total, multiply by three quarters, and that is your discount to the cent, with none of the assumptions about a typical workload that the 25% figure carries. What that figure requires is a bill a third made of cache reads, which is a much more specific claim than the announcement makes it sound.

The break-even did not move, and that is the point people are getting wrong

A 75% cut to cache reads sounds like it should change the decision about whether to cache at all. It does not, and the arithmetic is short enough to show. Caching a prefix costs one write plus a read per subsequent call; not caching it costs base input every time. On the five-minute cache, $12.50 of write plus $1.00 per read beat $10.00 per uncached call after 0.278 of a read, which rounds up to one. On Fable 5.1 the same comparison is $12.50 plus $0.25 per read, and the threshold is 0.256 of a read. Also one.

The one-hour cache behaves the same way: 1.111 reads on Fable 5, 1.026 on Fable 5.1, which is two in both cases. Anthropic's own documentation still states the general rule, that caching pays off after one cache read for the five-minute duration and two for the one-hour, and the rule survives the cut intact.

The reason is that the break-even is set by the write premium, not the read price. A five-minute write costs 1.25x input, so it has a 0.25x hole to dig out of, and it digs out of that hole on the first read whether a read costs a tenth or a fortieth. Anthropic left the write multipliers alone, so it left the break-even alone. Caching on Fable 5.1 is not worth starting sooner than it was on Fable 5. It is worth more once started, which is a different claim and a smaller one than most of this week's writeups have made.

A cache miss now costs fifty reads instead of twelve and a half

Here is the half of the change nobody is pricing. Writes stayed still while reads fell by three quarters, so the ratio between them moved by exactly the same factor. On Fable 5, rebuilding a five-minute cache cost 12.5 times what reading it cost. On Fable 5.1 it costs 50 times. Against the one-hour write the ratio went from 20 to 80.

Put a workload behind that. A 100,000-token cached prefix reads for $0.025 on Fable 5.1 and rewrites for $1.25. If something near the front of your prompt changes between calls, a timestamp in the system message, a tool list you build from an unordered map, a session id, you invalidate the prefix and pay the write again. That mistake used to cost you 12.5 reads. It now costs 50. The absolute penalty is identical, $1.25 either way, but relative to what you should have paid it got four times worse, and relative-to-correct is how anyone actually notices a caching bug in a bill.

The same shift changes what your caching bill is made of. Across a sustained hour on the five-minute cache, writes were 58.1% of the caching spend on Fable 5. On Fable 5.1 they are 84.7%. Cache reads have stopped being the thing worth optimising and cache writes have become it, which inverts the advice almost every prompt-caching guide gives, including ours from earlier this year.

The one-hour cache got twice as good without changing price

If writes now dominate, the lever that matters is how often you write, and Anthropic sells exactly one control for that: the one-hour cache at 2x input instead of 1.25x. It is 1.6 times the price of the five-minute write and it lasts twelve times as long. That trade was already good. It got substantially better on September 1, and not because its price changed, since its price did not change.

ConfigurationWritesReadsCost for the hour
Fable 5, 5-minute cache12108$25.80
Fable 5, 1-hour cache1119$13.90
Fable 5.1, 5-minute cache12108$17.70
Fable 5.1, 1-hour cache1119$4.975

A 100,000-token prefix, one call every thirty seconds for an hour, 120 calls. The five-minute cache expires and is rewritten twelve times; the one-hour cache is written once. On Fable 5 the one-hour option was 1.86 times cheaper. On Fable 5.1 it is 3.56 times cheaper, because the reads it buys in bulk got cheap while the write it charges for did not. Anyone running a long-lived agent on the default five-minute TTL is leaving a larger share on the table today than they were on Monday.

What forty turns of a coding agent actually bills

Abstract multipliers are easy to nod along to, so here is one concrete session. Forty turns. Each turn reads a 120,000-token cached prefix, adds 4,000 tokens of fresh input, and produces 2,500 tokens of output. The whole thing runs inside twenty minutes, so the five-minute cache is rewritten four times and the one-hour cache once.

Model and TTLWritesReadsFresh inputOutputSession
Fable 5, 5-minute$6.00$4.32$1.60$5.00$16.92
Fable 5, 1-hour$2.40$4.68$1.60$5.00$13.68
Fable 5.1, 5-minute$6.00$1.08$1.60$5.00$13.68
Fable 5.1, 1-hour$2.40$1.17$1.60$5.00$10.17
Opus 5, 5-minute$3.00$2.16$0.80$2.50$8.46
Opus 5, 1-hour$1.20$2.34$0.80$2.50$6.84

Holding the TTL constant, the new model saves 19.15% on the five-minute cache and 25.66% on the one-hour. Neither is 45%, because cache reads are only a quarter of this session's Fable 5 bill rather than the 60% that figure requires, and the check works: 25.53% of the bill times 0.75 is 19.15% exactly.

The row worth staring at is the middle pair. Fable 5.1 on the five-minute cache and Fable 5 on the one-hour cache both bill $13.68. For this workload, switching the TTL on the old model was worth precisely as much as the entire 75% price cut is on the default one. That is a coincidence of these particular parameters and not a general law, but it does put the size of the announcement in proportion against a configuration change that has been available all along and costs nothing to make.

Against Opus 5, the same session runs $8.46. Fable 5 was exactly twice that, since every line of Opus 5's card is exactly half of Fable 5's. Fable 5.1 narrows the gap to 1.617x. The flagship is still the expensive option; it is just less expensive than the clean 2x it used to be, and only on workloads that read a lot of cache.

On this one line the model ladder compresses by a factor of four

Anthropic's lineup spans a 10x range on base input, from Haiku 4.5 at $1.00 to Fable at $10.00. On cache reads that range has just been squeezed. Fable 5.1 costs 5.0 times Sonnet 5 on base input and 1.25 times Sonnet 5 on a cache hit. It costs 10 times Haiku 4.5 on base input and 2.5 times on a cache hit. Every ratio involving Fable 5.1 and a 0.1x model is a quarter of what the base prices imply.

For a workload dominated by cache reads, and long-running agents are exactly that, the practical distance between Anthropic's cheapest and most capable models has narrowed sharply on the input side. It has not narrowed at all on output, where Fable 5.1 is still $50.00 against Sonnet 5's $10.00 and Haiku 4.5's $5.00. Whether the flagship is now affordable for your agent therefore comes down to how talkative it is, not how much context you feed it. Our calculator takes cache reads and writes as separate inputs, which is the only way to see this.

Batch and US-only inference stack on top, in both directions

Anthropic states that the caching multipliers stack with other pricing modifiers, including the Batch API discount and data residency, so the 0.025x compounds. Batch halves everything, which puts a Fable 5.1 cache read at $0.125 per million against Fable 5's batch read of $0.50. That is the cheapest input token Anthropic has ever attached to a flagship model, and it is half of Opus 5's batched cache read of $0.25. Batch does not change any ratio, since it halves both sides; it just moves the whole ladder down and takes the 0.5x against Opus 5 with it.

The multiplier runs the other way too. Pinning inference to the United States with inference_geo set to us applies 1.1x to every token category including cache reads, so the read becomes $0.275. Google's regional Vertex endpoints carry their own 10% premium with the same effect. If you have a data residency requirement, your share of this price cut is 10% smaller than everybody else's, which is a small point except that it now applies to the only line that got cheaper.

Every venue shipped the new number on day one

Anthropic's announcement says Fable 5.1 is available today on all platforms, including Amazon Web Services, Google Cloud and Microsoft Azure, and the endpoint data backs it up: the Anthropic first-party, Bedrock, Vertex and Azure endpoints all carry the same per-token figures, including the $0.25 read. This is not how these rollouts usually go. Grok 4.6 took from August 18 to August 26 to reach its third cloud, and we spent a whole post on the ways those four cards ended up disagreeing.

The exception is regional routing. Vertex's European endpoint for Fable 5 lists $11.00 input and $1.10 on a cache read, a flat 1.1x of the global card, and the same premium should carry to 5.1. Anthropic documents the regional surcharge on Bedrock and Google Cloud for Sonnet 4.5, Haiku 4.5, Opus 4.5 and everything after them, so this is expected rather than surprising, but it means the $0.25 headline is a global-endpoint number and a European deployment pays $0.275 for it.

The other model on this multiplier is one you cannot buy

The 0.025x footnote names two models. Mythos 5.1 is the second, and Anthropic's pricing table marks it limited availability with a link to Project Glasswing, the vetted-partner programme that has gated every Mythos release. Its card is identical to Fable 5.1's down to the cent, and it scores 60.9% on Terminal-Bench 4.0 against Fable 5.1's 55.8%, so the cheapest cache reads Anthropic sells are attached to the model with the highest score and the shortest list of people allowed to call it.

That leaves the new multiplier applying to two rows out of seventeen, one of which is not generally available. It is a narrow exception for now. Whether 0.025x is the new standard or a one-generation promotion on the flagship is the question the footnote does not answer, and Anthropic has form in both directions: it cancelled a scheduled Sonnet 5 increase last month rather than let an introductory rate expire, and it has also let plenty of promotional pricing quietly become permanent.

What the price is attached to

Anthropic's published figures, from the launch post. Fable 5.1 against Fable 5 and Opus 5, so the comparison is against the model it replaces and the cheaper one it is competing with inside Anthropic's own lineup.

BenchmarkFable 5.1Fable 5Opus 5
Terminal-Bench-Science 0.152.6%24.7%29.0%
Terminal-Bench 4.055.8%42.0%52.3%
GDPval-AA v2185317231824
OSWorld 2.0, partial77.9%72.9%75.4%
OSWorld 2.0, strict41.7%36.1%39.6%
Humanity's Last Exam, no tools60.9%57.8%56.6%
Humanity's Last Exam, with tools65.0%63.8%63.6%
AutomationBench31.4%17.1%26.9%
CursorBench 3.2.073.4%70.5%70.0%

Anthropic's own table carries a fourth column for GPT-5.6 Sol, which we have dropped because it is populated on only five of these nine rows. On Terminal-Bench 4.0 the Fable 5.1 cell also carries a parenthetical 60.9% for Mythos 5.1, the higher of the two. Terminal-Bench-Science is the outlier of the set: 24.7% to 52.6% is a factor of 2.13, and no other row moves anything like that far. GDPval-AA v2 goes 1723 to 1853, passing Opus 5's 1824. All of it is vendor-reported and none of it has an independent replication two days after launch.

All prices read from Anthropic's pricing documentation on September 2, 2026. List prices only; Anthropic negotiates volume and enterprise discounting that no public page reflects.

One line changed and the footnote is where they said so

  • Anthropic: Claude platform pricing - The document this post turns on. Seventeen model rows with base input, both cache writes, cache hits and output, plus the footnote stating that cache hits and refreshes on Fable 5.1 and Mythos 5.1 are priced at 0.025x the base input price and that all other models use the standard 0.1x multiplier. Also the source of the batch table, the 1.1x data-residency multiplier, the stacking rule, and the sentence about caching paying off after one read on the five-minute cache and two on the one-hour. Read September 2, 2026
  • Anthropic: Introducing Claude Fable 5.1 and Claude Mythos 5.1 - September 1, 2026. The 75% and $0.25 per million figures, the estimated 25% for typical workloads and up to approximately 45% for highly agentic work, the availability sentence covering AWS, Google Cloud and Microsoft Azure, and every benchmark in the table above
  • OpenRouter: Claude Fable 5.1 endpoint detail - Pulled September 2, 2026. Four endpoints, Anthropic first-party, Bedrock, Vertex and Azure, all carrying $10.00 input, $50.00 output, $0.25 cache read, $12.50 and $20.00 cache writes, 1,000,000 context and 128,000 max completion. The companion endpoint list for Fable 5 supplies the $1.00 read and the Vertex Europe row at a flat 1.1x, which is where the regional figures in this post come from, since neither the AWS nor the Google pricing page renders its Claude table through our fetcher
  • LiteLLM: model prices and context window - The claude-fable-5-1 entry, added upstream within a day of launch, independently carries cache_read_input_token_cost of 2.5e-07 against Fable 5's 1e-06, along with prompt_cache_min_tokens of 512 and a provider_specific_entry of 1.1 for us. Its deprecation_date of 2027-09-01 matches the retirement commitment Anthropic publishes on its own models overview, not sooner than September 1, 2027, which is a year to the day from launch
  • Anthropic: models overview - The model id claude-fable-5-1, the 1M window, 128K max output, adaptive thinking always on at default effort high, a June 2026 knowledge cutoff, and the retirement date above. Also the footnote still asserting that prompt cache reads cost 10% of the base input price, on the same page and in the same table as Fable 5.1, and the line that now files Fable 5 under legacy models still available, one day after it was the flagship. The note that the current tokenizer arrived with Opus 4.7 is applied here to the whole current lineup, Fable 5.1 included
  • TokenCost: prompt caching pricing in 2026 - Our cross-provider comparison of caching mechanics and rates, which assumed a stable 0.1x on the Anthropic side throughout and now needs the exception noted
  • What we could not establish. Whether 0.025x is permanent or a launch promotion: Anthropic gives it no end date and no introductory label, which is how permanent prices usually look, but the footnote wording treats 0.1x as the standard and leaves the exception unexplained. We could not read the AWS Bedrock or Google Vertex Claude tables from their own pages, so every cloud figure here is OpenRouter's reading of them rather than a first-party one, and the European 1.1x for Fable 5.1 specifically is inferred from Fable 5's European row plus Anthropic's documented regional policy rather than observed. The tokenizer footnote on the pricing page covers Claude 4.7 and later models and Mythos Preview and never names the Fable line, though the models overview applies the current tokenizer, introduced with Opus 4.7, to the whole present lineup with Fable 5.1 in it; every price comparison in this post stays inside the Fable family, so nothing here turns on it either way. The 25% and 45% decode to 33.3% and 60% cache-read shares exactly, which is either a deliberate way of quoting the figures or a coincidence of rounding, and Anthropic does not say which workloads produced them. And the benchmark table is entirely vendor-reported two days after launch