Skip to main content
TokenCost logoTokenCost
IndustryAugust 1, 2026·11 min read

OpenAI and Anthropic both charge exactly double for speed. One of them tells you how much speed you get, and no benchmark on earth is allowed to check either claim.

Neither Fast mode is new, exactly. What finished happening in the last eight days is that both labs converted latency from something you commit capacity for into something you buy per request. Anthropic stopped selling Priority Tier commitments and left Opus 5 unsupported by them entirely. OpenAI renamed Priority Processing to Fast mode on July 30 and raised the speed it targets. The result is a rate card with four rungs instead of two, spanning 4x from Batch to Fast on the same model weights, and a new question every buyer now has to answer without data. Because here is the part nobody has written down: speed is the one axis of that card that no independent benchmark measures. Artificial Analysis is forbidden from measuring it by its own published methodology, which excludes priority queues. OpenRouter returns null for throughput on Opus 5 Fast. Both vendors say "up to 2.5x". Exactly one contractual number exists in the whole category, OpenAI's floor of 80 output tokens per second for Sol Fast, and against a measured standard-tier 63.5 it guarantees 1.26x for a price of 2x.

Backlit car speedometer in darkness, illustrating LLM API latency tier pricing

Photo by CHUTTERSNAP on Unsplash

The four things that actually differ

  • Both charge precisely double. Sol Fast $10 and $60, Opus 5 Fast $10 and $50, every line exactly 2x. The 2x rule only holds for the current generation, though: GPT-5.5 Fast charges two and a half times list.
  • OpenAI backs it with an SLA (99% of Sol Fast requests above 80 output tokens per second). Anthropic publishes no floor at all, and says the gain is output throughput, not time to first token.
  • Anthropic's Fast covers the full context window. OpenAI's refuses requests above 272K prompt tokens outright, which is exactly the workload that hurts most.
  • No independent measurement of either tier exists, and none can. The benchmark everyone cites excludes priority queues by written policy.

The rate card grew a dimension

For most of the last three years an API price was a single pair of numbers, and the only lever underneath it was Batch: wait up to 24 hours, pay half. That is still true, and all three major labs have converged on precisely 50% for it. What has changed is the other end. There is now a rung above list price, and you reach it by setting one field.

TierGPT-5.6 SolClaude Opus 5Gemini 3.1 Pro
Fast / Priority$10.00 / $60.00$10.00 / $50.00$3.60 / $21.60
Standard$5.00 / $30.00$5.00 / $25.00$2.00 / $12.00
Batch / Flex$2.50 / $15.00$2.50 / $12.50$1.00 / $6.00
Top-to-bottom span4.0x4.0x3.6x

All figures per million tokens, input then output. Gemini rows are the sub-200K-token bands; above 200K Google doubles input and applies its own step on output. The span row prices a 50K input, 10K output request at each tier and divides top by bottom, which is why it is not simply the output ratio. Sol runs $1.10 on Fast, $0.55 on standard and $0.275 on Batch. Opus 5 runs $1.00, $0.50 and $0.25.

Two of those columns are exactly 2x from standard to Fast on every single line, which is a deliberate simplicity and worth noticing, because OpenAI does not apply it consistently across its own catalogue. The 5.6 family is uniformly doubled: Sol $10 and $60, Terra $4 and $24, Luna $0.40 and $2.40, with cached input doubling too. GPT-5.5 Fast is $12.50 and $75.00 against a $5.00 and $30.00 standard card, so 150% on top rather than 100%. GPT-4.1 Fast adds only 75%. If you are budgeting off a remembered rule of thumb, the rule holds for the current generation and nothing either side of it.

Google is the outlier in the other direction, at 1.8x rather than 2x, and it is the only one of the three that documents what each rung is supposed to feel like: Priority in "seconds", Standard in "seconds to minutes", Flex targeting one to fifteen minutes, Batch up to 24 hours. Vague, but it is more than the other two offer about the middle of their ladders.

OpenAI sells a floor. Anthropic sells a ceiling.

Both companies use the identical marketing phrase, "up to 2.5x". Underneath it they are selling structurally different products, and the difference is the single most useful thing in this post.

OpenAI attaches a service level to every Fast row. Sol Fast carries 99.9% uptime and a latency SLA of 99% of requests above 80 output tokens per second. Terra Fast is 70, Luna Fast is 100. Those are commitments, not aspirations, and they are the only enforceable speed numbers anywhere in this market right now.

Read the Sol and Luna floors next to each other for a moment. Luna Fast bills output at $2.40 per million and promises 100 tokens per second. Sol Fast bills output at $60.00 and promises 80. You pay 25 times more for a guaranteed speed that is one fifth lower. That is not incoherent, it is what happens when a bigger model is genuinely harder to serve fast, but it does mean the price of Fast mode is not the price of speed. It is the price of the model, doubled, with a speed commitment attached that varies inversely with how much you paid.

Anthropic publishes no SLA, no absolute tokens-per-second figure, and no floor of any kind. What it publishes instead is a careful qualifier, stated twice in its own docs: speed benefits are focused on output tokens per second, not time to first token. For a chat UI, where perceived responsiveness is dominated by how long the user stares at an empty box, that qualifier removes most of the value. For a batch of long generations where total wall clock is what hurts, it removes none of it. Anthropic is being unusually honest here and it is easy to read past.

The rest of Anthropic's Fast mode is hedged the same way. It is a research preview, gated behind an account manager or a waitlist. It runs on the first-party API only, not Bedrock, not Vertex, not Microsoft Foundry. It cannot be combined with the Batch API at all. Prompt caching does work, with the usual multipliers applying on top of the doubled base, though switching speed mid-flight invalidates your cache because requests at different speeds do not share cached prefixes. If you have built around a long shared prefix, that last detail is the one that will cost you money unexpectedly.

Nobody is allowed to check the number

We went looking for an independent measurement of either Fast tier and came back with nothing. Not "we could not find one". There is a structural reason one cannot exist, and it is written into the methodology of the benchmark everyone quotes.

Artificial Analysis's published methodology requires that the traffic it measures be served on the same publicly available configuration any ordinary developer would receive, and it explicitly excludes priority queues. That rule is correct and it is the reason its numbers are trustworthy. It also means AA is barred, by design, from ever publishing a Fast-tier figure. Both Fast model pages 404. OpenRouter, the other place people check, returns null for both latency and throughput on anthropic/claude-opus-5-fast, and carries no Sol Fast slug at all.

So consider what a buyer can and cannot comparison-shop in 2026. Price is published, exactly, by everyone. Intelligence is measured to three significant figures by half a dozen independent labs. Speed, the entire product being sold on this new rung of the ladder, is a vendor claim with an "up to" in front of it, and the infrastructure that would normally hold vendors honest has ruled itself out of the room. You will see posts asserting that early independent benchmarks confirm the 2.5x. There are no early independent benchmarks. Those posts are laundering the press release.

What you are buyingSol FastOpus 5 Fast
Price vs standard2.0x2.0x
Vendor speed claimUp to 2.5xUp to 2.5x OTPS only
Contractual floor99% above 80 tok/sNone published
Uptime SLA99.9%None published
Independent measurementNone possibleNone possible
Long contextNot supportedFull window
Works with BatchMutually exclusiveNo

What thirty cents of latency actually buys

You can still make a decision without an independent benchmark, as long as you price the range honestly rather than the midpoint. Take a 10,000-token answer from Sol, which is a realistic length for a long agent turn or a generated document. On standard output rates that generation costs $0.30. On Fast it costs $0.60. Ignore the input leg for a second, since the speed you are buying applies to generation.

Artificial Analysis currently reads Sol on the OpenAI endpoint at 63.5 output tokens per second, so that answer takes about 157 seconds on standard. At OpenAI's contractual floor of 80 tokens per second it takes 125, saving 32 seconds. At the advertised 2.5x, meaning 158.75 tokens per second, it takes 63, saving 94 seconds. Same thirty cents, two very different purchases.

ScenarioThroughputTime for 10K outSavedCost per minute saved
Sol standard (measured)63.5 tok/s157.5 sbaselinebaseline
Sol Fast at the SLA floor80 tok/s125.0 s32.5 s$0.55
Sol Fast at the claimed 2.5x158.8 tok/s63.0 s94.5 s$0.19
Opus 5 Fast at the claimed 2.5x134.0 tok/s74.6 s111.9 s$0.13

The last row is the best case only, because Opus 5 has no floor. Its worst case is not 1.26x like Sol's, it is undefined. Anthropic could serve you at standard speed during a capacity crunch and you would have no contractual claim, which is precisely the risk OpenAI charges the same 2x to eliminate.

Now you have a number you can actually hold against your own business. Somewhere between 13 and 55 cents per minute of latency removed, per concurrent request. If a support agent is waiting on that generation and their loaded cost is a dollar a minute, Fast pays for itself several times over even at the pessimistic end. If the generation is a nightly report nobody reads until morning, you are lighting money on fire and Batch is sitting right there at half of standard. Most workloads are neither, which is the honest answer: the question is not whether Fast is worth it, it is which fraction of your traffic is latency-sensitive enough to route there, and almost nobody has that number instrumented.

The 272K trapdoor, and the input price that is already doubled

Every Fast row in OpenAI's table carries the label "excludes long context", with a footnote defining that as requests estimated above 272K prompt tokens. The guide is blunter: long context, fine-tuned models and embeddings are not supported. So the workload with the most painful wall-clock time, the enormous agent context that takes minutes to chew through, is the one workload that cannot buy its way out.

That is also good news in one narrow sense, and it corrects something we have seen asserted repeatedly: Fast mode and the long-context surcharge do not stack. They are mutually exclusive. Nobody is going to bill you $20 per million input tokens.

But put the two rate cards side by side and something odd falls out. Above 272K tokens, standard processing bills input at 2x and output at 1.5x for the whole request, which takes Sol to $10.00 and $45.00. Fast mode bills $10.00 and $60.00. The input prices are identical. A 300K-token prompt on standard Sol pays exactly the Fast-tier rate on every input token and receives none of the speed, no SLA, and no option to upgrade. It is the worst cell in the entire grid and you land in it by writing a long prompt.

Anthropic does the opposite, explicitly. Its docs state that Fast mode pricing applies across the full context window, including requests over 200K input tokens. If your product is a large-context coding agent, that single sentence is probably worth more than the SLA OpenAI publishes and Opus 5 does not. It is the clearest case in this whole comparison where the two vendors are not selling the same thing at the same price, and the naive read of "both charge 2x" hides it completely.

Both labs closed the door you used to buy speed through

The old way to buy latency was a capacity commitment: pay up front for reserved throughput, get better service. Within a week of each other, both labs shut that path on their flagship line, and they did it in opposite-looking ways that arrive at the same place.

OpenAI renamed. The changelog is unambiguous that Priority processing became Fast mode on July 30, and that Sol's Fast speed was raised at the same time. Nothing broke: service_tier accepts both "priority" and "fast", the project setting for Priority is equivalent to selecting Fast, and no deprecation notice exists for the old value. If you have priority in your request builder you can leave it there. What changed is the framing, from a capacity relationship to a per-request purchase.

Anthropic just closed the shop. Priority Tier capacity commitments are no longer available for purchase, and the service tiers doc lists Priority Tier as supported on all models except Mythos 5, Mythos Preview, Opus 5 and Sonnet 5. Read that exclusion list again: it is the entire current generation. On anything Anthropic shipped in the last two months, per-request Fast mode is not the best way to buy speed, it is the only way, and it is gated behind a waitlist.

For anyone holding an unexpired Anthropic commitment this is worth a calendar entry rather than a shrug. Your existing capacity does not cover the models you are most likely to want to migrate to, and the replacement mechanism costs double list with no floor attached. That is a real change to a real budget, announced in a documentation page rather than a blog post.

Six times the throughput for 1.85x the price, on identical weights

There is a control group for all of this, and it makes both frontier Fast tiers look expensive. The specialised accelerator hosts sell exactly one thing, speed on open-weight models, and they have to publish real throughput numbers because that is the entire pitch. Groq lists tokens per second directly in its rate card. Cerebras does the same.

Both serve gpt-oss-120b, the same weights, which makes the comparison clean in a way nothing else in this post is. Cerebras runs it at roughly 3,000 tokens per second for $0.35 and $0.75. Groq runs it at about 500 for $0.15 and $0.60. On a 10K input, 2K output request that is half a cent against about a quarter of a cent, so the premium is roughly 85% for something like six times the throughput. Call it 3.2 units of speed per unit of money.

Measure the frontier tiers the same way and OpenAI's best case comes to 1.25, while its guaranteed case comes to 0.63, which is the arithmetic way of saying that at the SLA floor you are paying more per unit of speed after upgrading than you were before. The accelerator market is somewhere between two and five times better on the specific thing being sold. Nobody should pretend this is a like-for-like swap, because it is not: you cannot buy Sol or Opus 5 from Cerebras at any price, and if your task genuinely needs a frontier model then the open-weight route is not on the table. But it does establish what speed costs when it is sold competitively by vendors who have to prove the number, and the answer is a lot less than double. We went through the same dynamic when ten hosts landed on nearly the same price for K3, and the pattern holds: where buyers can compare, margins compress.

Buy the floor, not the ceiling

Default to standard and treat Fast as a routing decision for a minority of traffic, not a global setting. The failure mode here is flipping service_tier at the client level and doubling the entire bill to speed up the 8% of requests a human is actually watching.

Before you enable anything, instrument two numbers you almost certainly do not have: what fraction of your requests have a human blocked on them, and what a minute of that person's waiting is worth. Fast mode costs between 13 and 55 cents per minute saved on Sol. If you cannot say whether that is cheap or expensive for your product, the answer is that you are not ready to buy it yet.

On OpenAI, prefer Fast where you have an SLA to inherit. The 80 tokens per second floor on Sol is a number you can put in your own contract, and that is a genuinely different product from a marketing claim. On Anthropic, be more careful: with no floor published, budget for the possibility that you paid double and got standard speed, and validate on your own traffic before you commit. Watch the time-to-first-token caveat if you are building a chat surface, because Anthropic has told you plainly that this is not what Fast mode improves.

And check the other direction first. The Batch rung is a flat 50% at all three labs, it is boring, it is unglamorous, and for most background workloads it is a larger and far more certain win than Fast mode is on the way up. A 4x span means the distance from Fast to Batch is bigger than most of the model-to-model price differences people agonise over. Run your own mix through the calculator at each tier before you decide which rung you belong on.

Sources

  • OpenAI: Fast mode - The published Fast rate table and SLAs: Sol $10.00 / $1.00 / $60.00 at 99% above 80 tok/s, Terra $4.00 / $24.00 at 70 tok/s, Luna $0.40 / $2.40 at 100 tok/s, all marked as excluding long context above 272K prompt tokens
  • OpenAI: Fast mode guide - "Up to 2.5x faster speeds and more consistent latency", the statement that long context, fine-tuned models and embeddings are unsupported, and confirmation that service_tier accepts both priority and fast
  • OpenAI: API changelog - The July 30, 2026 entry recording that Priority processing was renamed Fast mode and that Sol's Fast speed was raised to up to 2.5x standard
  • OpenAI: GPT-5.6 Sol model page - Standard $5.00 / $0.50 / $30.00, and the long-context rule pricing prompts above 272K input tokens at 2x input and 1.5x output for the full request
  • Anthropic: Fast mode - "Up to 2.5x higher output tokens per second", the twice-stated caveat that gains are OTPS and not TTFT, the Batch API exclusion, and the note that requests at different speeds do not share cached prefixes
  • Anthropic: Pricing - Opus 5 standard $5 / $25 with $0.50 cache hits and $2.50 / $12.50 batch, Fast at $10 / $50, and the statement that Fast pricing applies across the full context window including requests over 200K input tokens
  • Anthropic: Service tiers - Priority Tier capacity commitments are no longer available for purchase, and Priority Tier is unsupported on Opus 5, Sonnet 5, Mythos 5 and Mythos Preview
  • Artificial Analysis: Performance benchmarking methodology - The requirement that measured endpoints be the same publicly available configuration any developer receives, explicitly excluding priority queues, which is why no Fast-tier measurement exists
  • Google: Priority inference - The four-tier Priority, Standard, Flex and Batch ladder with latency targets, and Gemini 3.1 Pro Priority at $3.60 / $21.60 against $2.00 / $12.00 standard
  • Groq: Pricing - gpt-oss-120b at 500 tokens per second for $0.15 / $0.60, with throughput published directly in the rate card
  • Cerebras: Pricing - gpt-oss-120b at roughly 3,000 tokens per second for $0.35 / $0.75, the cleanest same-weights speed comparison available