Skip to main content
TokenCost logoTokenCost
ResearchMay 5, 2026·10 min read

Three weeks of Opus 4.7 bills are in. The tokenizer change costs an extra 25 to 37 percent in production.

Anthropic shipped Claude Opus 4.7 at $5 input, $25 output per million tokens, identical to Opus 4.6. The launch posts called it flat pricing. Three weeks later Finout, OpenRouter, CloudZero, and Simon Willison have published the actual numbers, and the tokenizer that quietly changed underneath the rate card is hauling bills up 25 to 37 percent on the workloads people run most.

Streams of orange and red text fragments flowing across a dark surface representing Opus 4.7 tokenizer cost inflation

Photo by Vishnu Mohanan on Unsplash

The receipts came in fast. By April 27, eleven days after launch, Finout and OpenRouter had production billing data on Opus 4.7. Simon Willison had run side-by-side token counts a week earlier. CloudZero published a workload model on April 21. Anthropic's own migration guide concedes the tokenizer multiplier in dry, deniable language: same input maps to roughly 1.0 to 1.35 times the tokens depending on content type. The community has now filled in where in that range your specific bill lands.

The rate card that did not change

Same dollar figures, top to bottom. Anthropic kept the standard tier, the cache discount, and the batch discount identical between Opus 4.6 and 4.7.

TierInput / 1MCached / 1MOutput / 1MNotes
Opus 4.6 Standard$5.00$0.50$25.00Old tokenizer, ~750k words / 1M tokens
Opus 4.7 Standard$5.00$0.50$25.00New tokenizer, ~555k words / 1M tokens
Opus 4.7 Batch$2.50-$12.5050% off, 24h SLA
Opus 4.7 Cache write (5min TTL)$6.25--1.25x premium for first write
Opus 4.7 Cache write (1h TTL)$10.00--2x premium, longer TTL

The text-density column is the part the launch coverage skipped. Anthropic's own model docs page admits it: 1M Opus 4.7 tokens equal roughly 555k words, where the same 1M tokens on 4.6 carried 750k. That is a 26 percent text-density loss baked directly into the official documentation, sitting behind a tooltip nobody reads.

What the new tokenizer actually does to the same bytes

Simon Willison's token comparisons were the cleanest first read: feed identical bytes to both tokenizers, count what comes out. The Tokenomics community ran a follow-up across more content types. The picture that emerged is not a single multiplier - it is a distribution.

Content typeOpus 4.6 tokensOpus 4.7 tokensMultiplier
System prompt (Willison sample)5,0397,3351.46x
Real Claude Code sessionbaseline+32.5%1.325x
Technical documentationbaseline+47%1.47x
CLAUDE.md filesbaseline+44.5%1.445x
30-page PDF (mixed prose)56,48260,9341.08x
Standard PNG image (682x318)3103141.01x

System prompt and PDF figures from Simon Willison, April 20, 2026. Claude Code, documentation, and CLAUDE.md figures aggregated by the Tokenomics Discord across roughly 200 sample files. Image figure from Willison's standard test PNG.

The pattern is what you would expect from a tokenizer optimized for a different distribution. Code, structured config, and Markdown with backticks expand the most because the new vocabulary breaks symbols differently. Plain prose barely moves. Images and audio do not move at all. If your workload is heavy on system prompts, CLAUDE.md files, or technical docs, you are sitting near the top of the inflation range.

What three weeks of production billing actually showed

OpenRouter sees billing across thousands of customers and bucketed Opus 4.7 against 4.6 by prompt size. The result is the closest thing to a public production audit we have right now. The cache absorbs much more than the launch posts implied.

Prompt sizeCost change vs 4.6What is going on
Under 2K tokens-1.6%Short prompts compress slightly better
2K to 10K tokens+27.2%The pain band - typical chat and RAG sit here
10K to 25K tokens+25.2%Agent loops and longer documents
128K and above+15.3%Cache hits absorb 93% of the extra tokens at $0.50/M

OpenRouter aggregated billing data, April 27, 2026. Cost change is measured at the session level, holding completed work constant.

The most expensive band is exactly where the bulk of agent traffic lives. A standard customer support workflow with retrieval, four to eight tool calls, and a wrapped reply lands in the 2K to 10K bucket. Coding agents on a single function change tend to fall in 10K to 25K. Both pay the full tokenizer tax.

The 128K-plus number is the part nobody on the launch coverage saw coming. Above 128K, the long preamble that drives the inflation - system prompt, big tool descriptions, CLAUDE.md, repeated context - tends to be cached. OpenRouter measured 93 percent of the tokenizer-induced extra tokens landing on cache hits at $0.50 per million, so the marginal cost increase shrinks. Long context is paradoxically the cheapest place for Opus 4.7's tokenizer to live.

One coding agent, one bill, two months

Finout's headline workload was a coding agent that ran $300/month on Opus 4.6 in March. Same scaffold, same task volume, same prompts. April 16 forward they switched to Opus 4.7 and watched the bill resettle.

Line itemMarch on 4.6April on 4.7Delta
Input tokens$120$162+35%
Output tokens$150$203+35%
Cache reads$30$40+33%
Total$300$405+$105 (+35%)

The delta is uniform across line items because the tokenizer change touches both sides equally. Inputs that were 100k tokens on 4.6 became 135k on 4.7 at the same $5 per million rate. Outputs that were 40k became 54k at the same $25. Cache hits scaled proportionally. Nothing on the price sheet moved. Everything on the token counter did.

CloudZero modeled an 80-turn agent session against the same workload assumption and landed at $7.86 to $8.76 versus $6.65 on Opus 4.6. That is a 20 to 32 percent jump on a per-session basis, which lines up with the OpenRouter aggregated number once you account for shorter sessions skewing the average down.

Three levers to claw the bill back

The pricing levers Anthropic kept flat are the only way to neutralize the tokenizer hike without leaving Opus 4.7. Stack them in the right order.

One. Cache aggressively. Cache hits at $0.50 per million absorb 90 percent of the input cost. OpenRouter's data shows that at 128K-plus, this single change drops the inflation from the 27 percent band down to 15 percent. Move CLAUDE.md, system prompts, and tool definitions into 1-hour TTL caches with the 2x write premium and recover the cost on the second call.

Two. Trim the system prompt. Willison's 5,039-to-7,335 example is the highest-leverage place to look. The new tokenizer is most punishing on dense Markdown, code fences, and long enum lists. Half the system prompts we have audited carry redundant context that compresses cleanly. A 30 percent system prompt trim usually gets the inflation back under 10 percent for chat workloads.

Three. Ship the slow path on Batch. Anthropic's 50 percent discount on Batch ($2.50 input, $12.50 output) more than offsets the worst-case tokenizer hit. If a job can wait 24 hours - data generation, evaluations, retroactive enrichment - Batch is the cheapest knob you have. The combination of Batch and cache reads can keep effective costs flat versus Opus 4.6 even on the highest-inflation content types.

When to leave Opus 4.7 entirely

Same token bill on May 5, 2026, across the most plausible substitutes:

ModelInput / 1MOutput / 1MEffective vs 4.7 (after tokenizer)SWE-Bench Pro
Claude Opus 4.7$5.00$25.00baseline64.3
Claude Sonnet 4.6$3.00$15.00~56% cheaper~55
GPT-5.5$5.00$30.00~10% more on output58.6
Gemini 3.1 Pro$2.00 / $4.00$12.00 / $18.00~55% cheaper at <200K~56
DeepSeek V4-Pro (promo)$0.435$0.87~95% cheaper~57
Kimi K2.6$0.60$2.50~88% cheaper58.6

Effective cost compares tokens-out-the-door for the same workload, applying the measured 1.30 average tokenizer multiplier to Opus 4.7 traffic. Other models use their stable tokenizers. SWE-Bench Pro figures from public leaderboards as of May 5, 2026. DeepSeek V4-Pro promo runs through May 5; regular pricing reverts to $1.74/$3.48 thereafter.

Sonnet 4.6 is the easiest substitution. Same provider, same SDK, same prompts work unchanged, and the older tokenizer means tokens-per-byte stays where you measured it. The benchmark gap to Opus 4.7 is real but narrower than the cost gap. For most customer-facing chat and RAG workloads, Sonnet 4.6 with caching is the new default.

Coding agents are where Opus 4.7 still earns its keep. SWE-Bench Pro at 64.3 beats every alternative on this list. The right framing is not whether to use Opus 4.7 - it is whether the marginal task is worth the now-25-to-37-percent premium over the baseline you measured a month ago. Bench against your own scaffold and only keep the high-value calls on Opus 4.7. Route everything else to Sonnet 4.6 or Kimi K2.6.

What this episode actually was

Anthropic shipped a price hike. They shipped it through the tokenizer instead of the rate card, which let the launch material claim flat pricing in good faith while every customer's billing chart shows a 25 to 37 percent step function on April 16. The disclosure in the migration guide is real but buried, the text-density tooltip is real but tooltip-shaped. Most teams got blindsided.

The right read on Opus 4.7 is not that it is a worse deal than Opus 4.6. It is that the deal you priced into your runway in March is gone. If your workload is short prompts under 2K, you got a microscopic discount. If your workload is long cached context above 128K, you got a 15 percent hike. Everyone in the middle - the chat apps, the RAG backends, the coding agents that live in 2K-25K prompts - just got a quiet 25-to-30-percent bill increase wearing a flat-pricing label.

Audit your token counts. Cache, batch, or trim until the math is back where you had it, or move to a substitute that did not change tokenizers. Do not pay 35 percent more for the same work because the launch post said the price was the same.

Sources