Claude Sonnet 5's $2 and $10 promo ends on August 31. Every line on the rate card moves up 50% at once, and there is no lever that softens it.
August 31 is 33 days out. After that, Anthropic's mid-tier model bills at $3 per million input tokens and $15 per million output, up from $2 and $10. This was never hidden: the expiry sat in the launch post on June 30 and has been a second dated row in the pricing docs ever since. What is worth an afternoon of planning is how evenly it lands. Caching and the Batch API both keep working exactly as well as they do now, which is precisely the problem, because their discounts are percentages and the percentages do not change. Below is the full before-and-after card, one real workload priced on both sides of the date, the tokenizer detail that makes "unchanged from Sonnet 4.6" misleading, and where $3/$15 leaves Sonnet 5 against everything else you could run.

Photo by Pablo Añón on Unsplash
Four dates on the calendar
- June 30, 2026
- Sonnet 5 ships. The launch post names $2/$10 as introductory and gives the expiry date in the same sentence, so nothing about September is a surprise.
- August 31, 2026
- Last day at $2 input and $10 output per million. Cache writes, cache hits and both Batch API lines are on the same promotional card and expire with it.
- September 1, 2026
- Everything moves to $3/$15 and its matching cache and batch rates. Sonnet 5 lands on the same card Sonnet 4.6 carries, while billing about 30% more tokens for identical text, and passes Opus 4.8 on cost per finished task.
- The gap in between
- Anthropic has published no rule for a Batch API job submitted on August 31 that completes on September 1. Batches can legitimately run 24 hours, so the boundary is real and undocumented.
The rate card, before and after
Anthropic ships this as two labelled rows in its pricing tables, one reading "through August 31, 2026" and one reading "starting September 1, 2026". Every figure in the second row is the first row times 1.5. That is unusual. Most repricing events move the headline rates and leave the discount tiers alone, or move output more than input to nudge you toward a particular usage shape. This one simply scales the whole card.
| Line item | Through Aug 31 | From Sept 1 | Change |
|---|---|---|---|
| Input | $2.00 | $3.00 | +50% |
| Output | $10.00 | $15.00 | +50% |
| Cache write, 5 min | $2.50 | $3.75 | +50% |
| Cache write, 1 hour | $4.00 | $6.00 | +50% |
| Cache hits and refreshes | $0.20 | $0.30 | +50% |
| Batch input | $1.00 | $1.50 | +50% |
| Batch output | $5.00 | $7.50 | +50% |
All figures are per million tokens. Two multipliers stack on top of whichever row applies to you. Setting inference_geo to US data residency adds 10% across every token category on Claude 4.6 and later, which turns the September card into $3.30 and $16.50. Regional and multi-region endpoints on Bedrock and Google Cloud carry the same 10% premium over global. Neither multiplier changes on September 1; they just apply to a bigger base.
One question the docs do not answer: what happens to a Batch API job submitted on August 31 that finishes on September 1. Batch jobs can legitimately run for 24 hours, so the boundary is real, and we could find no published rule on whether billing follows submission time or completion time. If you have large batches scheduled near the date, it is worth moving them a day earlier rather than finding out.
Anthropic never called this a price rise
The framing matters, because it tells you not to wait for a reversal. Anthropic's docs treat $3/$15 as the standard price and $2/$10 as the deviation. The what's-new page for Sonnet 5 opens by stating the model "is priced at $3 per million input tokens and $15 per million output tokens, unchanged from Claude Sonnet 4.6", and only then mentions the introductory rate; on the models overview the $2/$10 figure is a footnote. The launch post on June 30 put it plainly: the model was available "at an introductory price of $2 per million input tokens and $10 per million output tokens through August 31, 2026".
So this is a 63-day promotion lapsing on the schedule it was announced with, not a repricing anyone has to defend. Worth noting for anyone budgeting the rest of the Claude lineup: Sonnet 5 is the only model on the entire pricing page carrying two dated rows. Opus 5, Opus 4.8, Sonnet 4.6, Haiku 4.5 and Fable 5 all list a single set of rates with nothing scheduled. The only other date on the Claude calendar this summer is a retirement rather than a reprice, with Opus 4.1 tentatively due to be switched off on August 5, on the deprecations page rather than the pricing one.
One workload, priced on both sides of the date
Take a coding assistant burning 400M input and 40M output tokens a month, a shape with heavy prompt reuse. Three ways to buy it, all before and after. The cached row assumes a realistic split of 5% of input going to five-minute cache writes, 80% landing as cache hits and the remaining 15% billed as fresh input.
| How you buy it | Monthly, now | Monthly, Sept 1 | Extra per month |
|---|---|---|---|
| Straight API, no caching | $1,200 | $1,800 | +$600 |
| With prompt caching | $634 | $951 | +$317 |
| Batch API, no caching | $600 | $900 | +$300 |
The caching row saves 47.2% against the straight API, and it saves 47.2% in September too. Batch saves exactly half in both columns. That is the whole shape of this event: optimisation still pays, it just cannot pay more than it already does. Teams that already run a tight caching setup have no headroom left to claw back the difference, which is a slightly perverse outcome. The teams with the most room to absorb this are the ones who have not bothered optimising yet.
Swap in your own token mix on the calculator. It already stores both Sonnet 5 rate cards with the effective date attached, so it bills the promotional rate today and flips itself on September 1 rather than quietly going wrong.
Same card as Sonnet 4.6, about 30% more tokens on the meter
Here is the part that does not show up in any comparison table. Anthropic's pricing docs carry a note that Claude 4.7 and later models use a newer tokenizer which "produces approximately 30% more tokens for the same text". The what's-new page confirms Sonnet 5 is in that group and spells out the consequence: the same input text produces roughly 30% more tokens than on Sonnet 4.6, and "the cost of an equivalent request can differ from Claude Sonnet 4.6 even though per-token pricing is unchanged".
Put that next to the September rate card and the two statements interlock awkwardly. From September 1 Sonnet 5 charges $3/$15, and Sonnet 4.6 charges $3/$15. Identical prices. Feed both the same document and Sonnet 5 meters about 30% more tokens for it, so the invoice is about 30% higher for work that is textually identical. Today the promotional rate more than cancels that out: two-thirds of the price against 1.3 times the tokens leaves Sonnet 5 roughly 13% cheaper than Sonnet 4.6 on the same text. On September 1 that flips to roughly 30% dearer. The rate card will tell you nothing changed.
The same effect quietly shrinks the context window. Sonnet 5's 1M window holds noticeably less actual text than Sonnet 4.6's 1M window, because each token covers less of it. Anthropic is upfront about all of this in the docs; it just never puts the two facts in the same sentence. We went through the same exercise when the tokenizer first shipped with Opus 4.7, and the honest measurement then was that the multiplier varies by content, roughly 1.0x to 1.35x depending on how much code, markup and non-English text you send. Prose sits near the top of that range. If your workload is mostly source code, measure it before you assume 30%.
September 1 is also when Sonnet 5 passes Opus 4.8 on cost per task
Artificial Analysis runs every model through the same evaluation suite and publishes what it cost them, which is the only public number we know of that prices a finished job rather than a token. Their per-task figure for Sonnet 5 is about $2.29, and they are explicit that the number uses the standard $3/$15 rates rather than the promotion. Opus 4.8 came in around $1.99. So the September picture is a mid-tier model costing roughly 15% more per finished job than the flagship tier above it, while listing at 60% of its per-token price. Artificial Analysis put it plainly: Sonnet 5 "without promotional pricing will cost more per task than Opus 4.8", and they call it one of the most costly models to run, behind only Fable 5.
Right now the promotion hides that. At $2/$10 the same work lands nearer $1.53 per task, roughly 23% below Opus 4.8, which is exactly the bargain Sonnet 5 has looked like since June. The crossover is the date, not the model. And the mechanism is worth understanding because it is not going away: Sonnet 5 burns roughly 40% more output tokens per task than Sonnet 4.6 and takes around three times as many agentic turns on the knowledge-work evaluations. Adaptive thinking is on by default, and reasoning tokens bill at the output rate.
| Model | Index score | Cost to run the suite | List price |
|---|---|---|---|
| Kimi K3 | 57 | $2,437 | $3 / $15 |
| GPT-5.6 Terra | 55 | $2,010 | $2.50 / $15 |
| Claude Sonnet 5 | 53 | $4,010 | $2 / $10 promo |
| Gemini 3.6 Flash | 50 | $727 | $1.50 / $7.50 |
Sonnet 5 carries the second-cheapest rate card in that table and is comfortably the most expensive model in it to actually run: five and a half times Gemini 3.6 Flash for three more points, and 65% more than Kimi K3 for four points fewer. Note the basis before extrapolating. At 300M output tokens, $15 per million would exceed $4,010 on the output side alone, so that suite cost is measured at the promotional rate. Scale it by the announced 1.5x and September comes in near $6,015. Running the same evaluations on GPT-5.6 Terra would leave about $4,000 of that unspent, and on Gemini 3.6 Flash you would keep roughly 88 cents of every dollar. Cost per token has rarely been a worse predictor of cost per job than it is here.
Where $3 and $15 leaves Sonnet 5 in the market
Today Sonnet 5 undercuts GPT-5.6 Terra on input. In September it goes above it. Same 400M input and 40M output workload as before, run against everything in and around that bracket at published list rates.
| Model | Input / 1M | Output / 1M | Monthly bill |
|---|---|---|---|
| Claude Opus 5 | $5.00 | $25.00 | $3,000 |
| Claude Sonnet 5, from Sept 1 | $3.00 | $15.00 | $1,800 |
| Kimi K3 | $3.00 | $15.00 | $1,800 |
| GPT-5.6 Terra | $2.50 | $15.00 | $1,600 |
| Claude Sonnet 5, today | $2.00 | $10.00 | $1,200 |
| Grok 4.5, under 200K prompt | $2.00 | $6.00 | $1,040 |
| Gemini 3.6 Flash | $1.50 | $7.50 | $900 |
| GLM-5.2 | $1.40 | $4.40 | $736 |
| GPT-5.6 Luna | $1.00 | $6.00 | $640 |
| Claude Haiku 4.5 | $1.00 | $5.00 | $600 |
| DeepSeek-V4-Pro | $0.435 | $0.87 | $209 |
The coincidence worth staring at is Kimi K3. Moonshot lists $3.00 input, $15.00 output and $0.30 for cached input, which is post-September Sonnet 5 to the cent on all three numbers, with a 1M context window on both sides. It also scores four points higher on the Intelligence Index and cost Artificial Analysis 39% less to evaluate. We wrote up K3's rate card when Moonshot tripled its own prices to get there, and the irony is that a rate card which looked expensive for a Chinese lab in July looks like a straight price match with Anthropic in September.
None of this makes Sonnet 5 a bad model. It makes it a model whose price advantage was temporary and whose token consumption is high, which is a different problem from being overpriced. If you picked Sonnet 5 in July because it undercut Terra, that specific reason expires with the promotion.
What we would do with the 33 days
Start with an actual number rather than a multiplier. Pull last month's token counts out of the console, split input into cached and uncached, and multiply by 1.5. That is your September bill, and for most teams it is a smaller shock than the percentage sounds, because caching is already doing most of the work. A team spending $634 a month is looking at $317 more, not $600.
Then run a routing test rather than a model swap. The interesting comparison is not Sonnet 5 against Sonnet 5, it is which slice of your traffic actually needs it. Haiku 4.5 at $1/$5 handles classification, extraction and short-form generation at a third of the September rate, and it stays inside the Anthropic API so nothing else in your stack moves. Most of the teams we see running everything through the mid tier are doing it out of convenience.
Measure your own tokenizer multiplier too, before you trust the 30% figure in either direction. Send a representative sample through Sonnet 5 and Sonnet 4.6 and compare the reported input token counts. Code-heavy workloads sit lower in the range; English prose sits near the top. If you land at 1.1x rather than 1.3x, the Sonnet 4.6 comparison looks quite different, and Sonnet 4.6 is not going anywhere.
What we would not do is treat this as an emergency. The increase is 50% on a mid-tier model, disclosed nine weeks in advance, on a rate card that returns to what Sonnet 4, 4.5 and 4.6 have all cost. The thing genuinely worth acting on is the token consumption, not the price per token, and that has been true since June. Every current rate card is here if you want to see the whole board before you decide.
Sources
- Anthropic: Claude pricing - Both dated Sonnet 5 rows, cache and batch rates, the data-residency multiplier, and the tokenizer note
- Anthropic: Introducing Claude Sonnet 5 - Launch post, June 30 2026, stating the $2/$10 introductory price through August 31
- Anthropic: What's new in Sonnet 5 - The "unchanged from Claude Sonnet 4.6" framing and the new-tokenizer section
- Artificial Analysis: Claude Sonnet 5 - Intelligence Index 53, $4,010 to evaluate, 300M output tokens, cost per task against Opus 4.8
- OpenAI: API pricing - GPT-5.6 Terra at $2.50/$15 and Luna at $1/$6
- Google: Gemini API pricing - Gemini 3.6 Flash at $1.50/$7.50
- Moonshot: Kimi K3 pricing - $3.00 input, $15.00 output, $0.30 cached input
- xAI: Models and pricing - Grok 4.5 at $2/$6 below 200K prompt tokens, $4/$12 above
- Z.ai: GLM pricing - GLM-5.2 at $1.40/$4.40
- DeepSeek: API pricing - V4-Pro rates used in the comparison table