Claude Haiku 5.5 bills GPT-6 Luna's $0.10 and $0.50 to the cent. One token past 100,000 the whole request costs five times as much, and Luna doesn't move until 272,000.
Anthropic took 90% off Haiku on Wednesday and copied OpenAI's cheapest rate card down to the cache write. The copy stops at 100,000 tokens, and so does most of the bargain.

Photo by Tim Mossholder on Unsplash
- Under 100K prompt tokens, Haiku 5.5 and GPT-6 Luna bill identically, and Haiku leads every row of Anthropic's benchmark table.
- Between 100K and 272K, Luna costs 80% less for the same request. Past 272K it's about 60% less.
- Haiku thinks long at max effort, roughly three times Luna's output tokens per task, and that can matter more than the rate card.
- Leaving Haiku 4.5? Temperature, prefill and budget_tokens now return errors.
Two rate cards, one for each side of 100K
Haiku 5.5 went live on October 7 as claude-haiku-5-5. Haiku 4.5 charged $1 and $5, so the short-prompt price is a tenth of what it was. Haiku is now the only Claude model from 4.6 onward that doesn't get flat pricing across its full 1M window.
| Per 1M tokens | Prompt up to 100K | Prompt over 100K |
|---|---|---|
| Input | $0.10 | $0.50 |
| Output | $0.50 | $2.50 |
| Cache write, 5 min | $0.125 | $0.625 |
| Cache write, 1 hour | $0.20 | $1.00 |
| Cache read | $0.01 | $0.05 |
| Batch input / output | $0.05 / $0.25 | $0.25 / $1.25 |
Anthropic pricing page, read October 9, 2026. US-only inference (inference_geo: "us") is 1.1x on every line. No Fast mode for Haiku.
Read the small print on that second column. The prompt length counts every input token, cache reads and writes included, and once it passes 100,000 the whole request moves up. Output too. Anthropic says the higher rate applies even when part of the prompt is a cache hit, so a big cached system prompt doesn't get you under the line.
One more change people will miss: Haiku keeps the old 0.1x cache read. Opus 5.5 and Sonnet 5.5 bill cache hits at 0.05x. At $0.01 it hardly matters, but past 100K the read is $0.05, half of Sonnet 5.5's $0.10.
Where Luna pulls away
GPT-6 Luna charges $0.10, $0.01 cached, $0.125 cache write and $0.50 out, identical to Haiku. Its threshold is 272,000 tokens, and past it Luna goes to $0.20 and $0.75, not 5x. Here is one request at each prompt size, with 2,000 output tokens:
Our arithmetic at standard list prices, no caching, no batch. Luna's max input is 922K tokens, so 900K is about as far as the comparison goes.
At 100,000 tokens both requests cost $0.011. Add one token and Haiku's costs $0.055. From there to 272K the gap is a flat 5x, which is the zone most RAG and document jobs live in. After 272K Luna doubles its input and the gap settles at about 2.5x.
We think this is the most important line in the release, and launch headlines skipped it in favor of "matches Luna." It matches Luna on chat turns, classification and short agent steps. It does not match Luna on long documents.
The tokenizer moves the line to about 77K
Haiku 5.5 uses the tokenizer Anthropic introduced with Opus 4.7. Anthropic says it turns the same text into about 30% more tokens than Haiku 4.5 did. So a prompt that counted 77,000 tokens on Haiku 4.5 counts roughly 100,000 on Haiku 5.5. If your logs say your prompts peak around 80K, assume some of them will now cross.
It still comes out cheaper than the old model. A 150K-token Haiku 4.5 request with 2K output cost $0.16. The same text on 5.5 is about 195K tokens in and 2.6K out, all at the high rate: $0.104. That's 35% less, not 90%. Anthropic's own estimate for typical workloads is about 75% less. On list price alone, the long-context tier is half of Haiku 4.5's card.
You can check your own prompts with the token counter, then add 30% for Claude's newer tokenizer.
Anthropic's scores, and what they cost to get
The launch page compares Haiku 5.5 with Haiku 4.5, GPT-6 Luna and Sonnet 5.5. No Gemini model is on it.
| Benchmark | Haiku 5.5 | GPT-6 Luna | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 | 1620 | 1437 | 735 | 1840 |
| OSWorld 2.1 | 72.4% | 48.9% | 15.7% | 83.9% |
| Terminal-Bench 4.0 | 39.2% | 16.4% | 0.0% | 70.6% |
| FrontierCode 1.1 | 46.4% | 42.4% | - | 52.1% |
| Chartography (no tools) | 46.4% | 29.1% | 6.4% | 61.6% |
Anthropic's Haiku 5.5 announcement, October 7, 2026. Vendor-run.
OSWorld is the eye-catcher: 72.4% against Luna's 48.9%, up from 15.7% on the previous Haiku. If you run computer-use agents on a budget, this is the model to try first.
Independent numbers agree on the ranking and add a catch. Artificial Analysis scores Haiku 5.5 at max effort 43 on its Intelligence Index, against 38 for Luna. To get there it used about 162,000 output tokens per task, roughly three times Luna's 50,000. At the same $0.50 per million that is about $0.08 of output per task against $0.025. AA's provisional all-in figure is $0.21 a task for Haiku against $0.07 for Luna, so five points cost three times as much. AA flags that its Haiku number doesn't yet include the over-100K step-up, so if anything it's low.
Before you switch
Haiku 5.5 is not a drop-in swap for 4.5. Adaptive thinking is on by default at medium effort, and manual budget_tokens returns an error. So do assistant prefill and any non-default temperature, top_p or top_k. If your 4.5 code sets temperature to 0 for extraction, it will start failing with 400s.
Thinking is the cost lever. You can switch it off at high effort or below, and for classification and routing that's where we'd start. The 162K-tokens-per-task figure is at max effort; at the default you won't see anything like it, but you will see more output than Haiku 4.5 ever produced.
On Haiku 4.5 itself: Anthropic still lists it as active, with retirement "not sooner than" October 15, 2026. No deprecation notice has been posted and Anthropic gives at least 60 days, so nothing stops working next week. Haiku 5.5 runs on the Claude API, Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Bedrock and Vertex hadn't listed a price when we checked.
Our pick by prompt size
Under 100K tokens, we'd default to Haiku 5.5. Same price as Luna, ahead on every row of Anthropic's table, and AA measures it at about 240 output tokens a second. Keep effort low or thinking off and the output habit stays in check.
Between 100K and 272K, use Luna unless you've tested both and Haiku's quality gap is worth 5x. Above 272K it's closer, 2.5x, but still Luna on price. If you need Claude for a long-context job, Sonnet 5.5 at a flat $2 and $10 is only 4x Haiku's long rate and scores far higher.
To price your own mix, put it in the calculator with both models selected.
Sources
- Anthropic: pricing - both Haiku 5.5 tiers, cache, batch, data residency
- Anthropic: Haiku 5.5 overview and what's new - release date, context, tokenizer, breaking changes
- Anthropic: Claude Haiku 5.5 announcement - benchmark table, 75% estimate
- Anthropic: model deprecations - Haiku 4.5 status
- OpenAI: API pricing - GPT-6 Luna, both tiers
- Artificial Analysis: Haiku 5.5 - Intelligence Index, output tokens, speed