Skip to main content
TokenCost logoTokenCost
ComparisonJuly 15, 2026·8 min read

Claude runs five times apart from its cheapest tier to its priciest. The capability gap is nowhere near that wide, so most teams overpay by picking the wrong Claude.

Haiku 4.5 lists at $1/$5, Sonnet 5 at $3/$15, Opus 4.8 at $5/$25. Sticker to sticker that is a clean 5x. Priced by the actual job it stops being clean: the model that clears your workload is often two tiers below the one you defaulted to. We ran all three across four real workloads to show where each Claude earns its rate.

Abstract purple waves on a dark background representing Claude API pricing tiers

Photo by Mirella Callage on Unsplash

The short version

Haiku 4.5 handles anything well-scoped, high-volume, or latency-bound for a fifth of Opus money. Sonnet 5 is the workhorse, roughly 40% under Opus while landing a hair behind it on the benchmarks that matter, and it brings a 1M context window Haiku does not have. Opus 4.8 is worth its premium only on the hardest coding and reasoning, where a few accuracy points cover the cost. Two levers reshuffle all of this: prompt caching, which can drag an input-heavy Sonnet month down near uncached Haiku, and the fact that a routine job may not belong on Claude at all. Sonnet 5 also runs $2/$10 on introductory pricing through August 31 before the standard $3/$15 kicks in.

The three tiers, all-in

ModelInputOutputCache readBatch (in/out)Context
Haiku 4.5$1$5$0.10$0.50 / $2.50200K
Sonnet 5 (standard)$3$15$0.30$1.50 / $7.501M
Opus 4.8$5$25$0.50$2.50 / $12.501M

Per million tokens, from Anthropic's pricing docs. Sonnet 5 is $2/$10 on introductory pricing through August 31, 2026, with cache reads at $0.20 and batch at $1/$5. Opus 4.8 also offers a Fast mode tier at $10/$50 for 2.5x speed. Cache writes run higher than the base input rate: about 1.25x for a 5-minute window and 2x for one hour.

Five times on price is not five times on capability

The pricing ladder is evenly spaced, which quietly invites people to assume the capability ladder matches it. It does not. On the coding scores Anthropic reports, Opus 4.8 posts SWE-bench Verified 88.6 and SWE-bench Pro 69.2; Sonnet 5 sits at 85.2 and 63.2, a few points behind for roughly 1.7 times the money. Independent trackers like Artificial Analysis tell the same story: the three Claude tiers cluster far closer on measured quality than their prices do. Haiku 4.5 gives up more ground on hard reasoning, but on the bounded tasks it is built for, the gap you would notice in production is small and the gap on the invoice is a factor of five.

So the question is never "which Claude is best" in the abstract. It is which one your specific traffic can drop to without the output getting worse in a way you or your users would actually feel. For a lot of production traffic, that floor is lower than the tier people reach for by habit.

Priced by the job

Sticker rates are abstract until you attach them to a shape of work. Here are four workloads a real product tends to run, priced monthly on each tier at list rates with no caching. The last row uses the Batch API, where every tier is halved.

Monthly workloadHaiku 4.5Sonnet 5Opus 4.8
Support chat, 2M in / 0.5M out$4.50$13.50$22.50
RAG Q&A, 20M in / 2M out$30$90$150
Coding agent, 8M in / 6M out$38$114$190
Batch extraction, 50M in / 5M out$37.50$112.50$187.50

Sonnet 5 at standard $3/$15. On introductory pricing through August 31 the support month is $9 and the coding month is $76. Batch row applies the 50% Batch API discount to every tier. All figures rounded, list rates, no prompt caching.

The ratio barely moves across workloads, which is the useful part: whatever the shape, Sonnet costs about three times Haiku and Opus about 1.7 times Sonnet. That means the decision is not workload math, it is a capability call. If Haiku answers your support tickets acceptably, you are paying three to five times more than you need every time you route those tickets to Sonnet or Opus out of caution. The coding-agent row is where the premium can be worth it, because that is where the extra points show up in fewer failed runs.

Caching can quietly demote a tier

Everything above assumes you pay full input rate on every token, every call. A lot of real workloads do not work that way. If your requests share a stable prefix, a system prompt, a tool schema, retrieved documents, a repo snapshot, prompt caching bills that prefix at roughly a tenth of the input rate on reads: $0.30 per million on Sonnet 5 instead of $3, $0.50 on Opus instead of $5.

Take the RAG month. Say 18M of that 20M input is a stable document prefix you read against repeatedly, with 2M of fresh queries and 2M of output. Uncached, Sonnet 5 runs $90. Cached, the prefix reads at $0.30 instead of $3, and the bill drops to about $41 before you account for the one-time write. That lands a cached Sonnet 5 within shouting distance of an uncached Haiku 4.5 at $30, for a stronger model on the same job. Caching does not change the sticker ladder, but on input-heavy, repetitive traffic it can flatten it enough that the mid tier stops looking expensive.

The catch is that cache writes cost more than base input, so caching only pays when the prefix is reused enough to amortize the write. Chat with no shared context gets nothing from it. A retrieval or agent loop hammering the same prefix hundreds of times gets almost all of it.

When the answer is not a Claude at all

Haiku 4.5 is the cheapest Claude, but it is not the cheapest capable model. For routine classification, extraction, and retrieval answering, models outside Anthropic's lineup sit well below it. On that same 20M-in, 2M-out RAG month, Gemini 3.5 Flash runs about $48 and DeepSeek V4 Flash around $3.40, against Haiku's $30. Flash is close to Haiku; DeepSeek is an order of magnitude under it.

The reason to stay on Claude for a job is a Claude-specific strength: the 1M context on Sonnet 5 and Opus, the agentic and coding scores, or a toolchain already wired to the Anthropic API. When none of those apply and the task is genuinely routine, the cheapest Claude is still leaving money on the table. Put your real token mix into a side-by-side comparison before you assume the answer has to wear a Claude badge.

Which Claude, for which job

Haiku 4.5 is the default for high-volume, latency-sensitive, well-bounded work: support turns, routing, tagging, extraction, first-pass drafts. It clears most of what teams route upward on reflex, at $1/$5 and a 200K window. Start here and promote only what visibly needs it.

Sonnet 5 is the workhorse for anything where answer quality tracks model strength: real coding, multi-step agents, analysis over long documents. It is close enough to Opus on most benchmarks to be the sensible ceiling for the majority of production traffic, and the 1M context plus introductory $2/$10 through August make it easy to justify right now.

Opus 4.8 is the specialist. Keep it for the hardest coding, agentic, and math tasks where its 88.6 SWE-bench Verified and top-of-lineup reasoning turn into fewer retries and better final answers, or use Fast mode when latency matters more than the bill. Running your whole pipeline on Opus is how you end up paying five times over for work Haiku would have finished.

Sources