At 07:00 UTC tomorrow the cheapest model we can price goes up 5.00x, and the number that replaces it is not new. It is the card Inception has charged for Mercury 2 since March, with a fifth off the input line and nothing off the output line.
Mercury 2.5 has been selling since August 31 at $0.04 and $0.15 per million tokens. On September 8 it becomes $0.20 and $0.75. The discount is a flat 80% struck against every line at once, which makes the reversion exactly 5.00x for everybody: no prompt shape, no cache-hit rate and no reasoning setting moves it a cent either way. The countdown itself is published in exactly one place, and that place is not Inception.

Photo by Jason Leung on Unsplash
What changes at 07:00
- Input goes $0.04 to $0.20. Output goes $0.15 to $0.75. Cached reads go $0.004 to $0.02. Three lines, one multiplier, 5.00x on each. Most price changes we cover reward you for a particular prompt mix or punish you for another. This one does neither, which is unusual enough to be worth saying plainly.
- The price that lands is Mercury 2's price. Inception has charged $0.25 and $0.75 for Mercury 2 since March. Mercury 2.5's list card is $0.20 and $0.75. Output is identical to the cent. Input and cache reads are 20% lower. A generation later, the output line has not moved.
- The clock is on the distributor's page, not the maker's. OpenRouter's banner reads "Limited-time 80% discount via Inception through September 8, 2026 at 07:00 UTC". Inception's own site lists Mercury 2.5 at $0.20 and $0.75 with no discount and no date, its docs do not list the model at all, and its own API does not serve it.
- A second discount expires 33 hours later. Z.AI's 50% off GLM-5.3-Flash runs out at 16:00 UTC on September 9. Between now and Wednesday afternoon the cheapest model in our comparison basket changes hands twice without a single model changing.
- There are no benchmarks. Not from Inception, not from Artificial Analysis, not from OpenRouter. The only quality claim anywhere is "a 10+ point jump in intelligence over Mercury 2", and it does not name the index.
- Speed does not make this cheaper per hour. It makes it more expensive per hour. Mercury 2.5 and Mercury 2 charge the same $0.75 for a million output tokens and Mercury 2.5 emits about twice as many per second, so an hour of continuous generation costs roughly $1.27 against $0.63.
One multiplier, applied to everything
OpenRouter's endpoint data for Mercury 2.5 carries a literal discount: 0.8 field alongside the prices. It is not a marketing description, it is the arithmetic the platform is applying, and you can see it resolve: every promotional figure is exactly one fifth of the corresponding list figure. Mercury 2, on the same API, returns discount: 0.
| Line | Today | From Sep 8 | Change |
|---|---|---|---|
| Input, per 1M | $0.04 | $0.20 | 5.00x |
| Output, per 1M | $0.15 | $0.75 | 5.00x |
| Cached input read, per 1M | $0.004 | $0.02 | 5.00x |
| Cache write | not priced | not priced | - |
| Batch | none published | none published | - |
| Context-length tier | none, flat to 260K | none, flat to 260K | - |
The practical consequence is that there is nothing to optimise around. When Google doubles Gemini Flash on January 1, caching still saves you 90% and batch still saves 50%, so a team with a high cache-hit rate absorbs the rise better than one without. Mercury 2.5 has no batch tier, no cache-write price, no context cliff and a cache read that rises by the same 5.00x as everything else. Whatever your workload looks like, tomorrow it costs five times what it costs today. We rarely get to write a sentence that clean.
The new price is the old price
"Mercury 2.5 gets 5x more expensive" is going to be the headline this week and it is technically true and slightly misleading. What actually happens on Tuesday morning is that a promotion ends and Inception's ordinary card takes over, and that card is very close to the one it has been publishing all along.
| Line | Mercury 2 | Mercury 2.5 list | Ratio |
|---|---|---|---|
| Input, per 1M | $0.25 | $0.20 | 0.80x |
| Output, per 1M | $0.75 | $0.75 | 1.00x |
| Cached input read, per 1M | $0.025 | $0.02 | 0.80x |
| Context window | 128,000 | 260,000 | 2.03x |
| Max output tokens | 50,000 | 65,536 | 1.31x |
Output has not moved in six months. That is the line that dominates the bill on a generative workload, and it is the line Inception left alone across a version bump that doubled the context window and, by its own account, added ten-plus points of intelligence. Read generously, the company held its headline number flat through a capability upgrade, which is the same thing Anthropic did across three Opus releases. Read plainly, the cheap number you saw last week was a launch promotion and the real product has always cost about what the last one cost.
If you are on Mercury 2 today, the reversion is good news. Moving to 2.5 after Tuesday drops your input rate 20%, doubles your context window and, on the only speed measurement that compares the two the same way, roughly doubles throughput, all for an output price that does not change. That migration is worth doing on September 8 and it was worth doing on September 1.
Where the countdown is published
This is the part that made us want to write the post. Four Inception-controlled surfaces exist, and none of them tells you that you are currently paying a fifth of list.
- Inception's models page lists Mercury 2.5 (Preview) at "Input $0.20 / 1M Tokens, Cached Input $0.02 / 1M Tokens, Output $0.75 / 1M Tokens". No struck-through price, no discount, no date. Read it today and you would conclude the model costs five times what you are actually being charged.
- Inception's developer docs do not list Mercury 2.5. The pricing table there has two rows, Mercury 2 and Mercury Edit 2.
- Inception's own API does not serve Mercury 2.5. A call to its models endpoint returns one object,
mercury-2. The model is genuinely OpenRouter-exclusive at the moment. - Inception's blog has not announced it. The newest entry is dated August 11, three weeks before Mercury 2.5 appeared.
So the only two numbers a developer can act on live on somebody else's website: the price you are paying, and the hour it stops. We are not treating that as a scandal. Preview models get launched sideways all the time and OpenRouter is a legitimate primary source here, since it is the merchant of record and the discount is visible in its API as structured data rather than marketing copy. But it does mean that anyone who priced this model by reading the vendor's own site has been overestimating their bill by 5x for a week, and anyone who priced it by reading OpenRouter is about to be surprised in the other direction.
What it costs against everything else
One basket, applied identically: a million input tokens and 250,000 output tokens, the rough four-to-one shape of a search agent or a coding subagent, which is what Inception says this model is for. Every price below comes from the vendor's own page or, for Mercury 2.5, from the page it is sold on.
| Model | In / out per 1M | Basket cost | vs today |
|---|---|---|---|
| Mercury 2.5, discounted | $0.04 / $0.15 | $0.0775 | 1.00x |
| GLM-5.3-Flash, discounted | $0.075 / $0.25 | $0.1375 | 1.77x |
| Qwen3.8-Flash | $0.113 / $0.382 | $0.2085 | 2.69x |
| GLM-5.3-Flash, list | $0.15 / $0.50 | $0.2750 | 3.55x |
| DeepSeek V4 Flash, off-peak | $0.22 / $0.66 | $0.3850 | 4.97x |
| Mercury 2.5, list | $0.20 / $0.75 | $0.3875 | 5.00x |
| Mercury 2 | $0.25 / $0.75 | $0.4375 | 5.65x |
| GPT-5.6 Luna, short context | $0.20 / $1.20 | $0.5000 | 6.45x |
| GPT-5.4 nano | $0.20 / $1.25 | $0.5125 | 6.61x |
| Gemini 3.1 Flash-Lite | $0.25 / $1.50 | $0.6250 | 8.06x |
| DeepSeek V4 Flash, peak | $0.44 / $1.32 | $0.7700 | 9.94x |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | $0.9250 | 11.94x |
| Gemini 3.8 Flash | $0.75 / $3.75 | $1.6875 | 21.77x |
| Claude Haiku 4.5 | $1.00 / $5.00 | $2.2500 | 29.03x |
Today Mercury 2.5 is first by a wide margin. Tomorrow it is fourth among distinct models, behind GLM-5.3-Flash and Qwen3.8-Flash and, by a quarter of a cent, behind DeepSeek V4 Flash on its off-peak rate. It stays fourth after GLM's own discount lapses on Wednesday, which is the sort of thing you only notice if you rank the whole table rather than the headline. That $0.3850 against $0.3875 is our favourite kind of coincidence: two models that could hardly be less alike, one a diffusion decoder sold through a single marketplace and one a mixture-of-experts sold direct with a time-of-day surcharge, landing within 0.65% of each other on the same basket. Neither of them planned it.
Scaled to a month that reads a billion input tokens and writes 250 million, the same arithmetic gives $77.50 today and $387.50 from Tuesday, a difference of $310.00. Small enough that nobody reorganises an architecture over it, large enough that the finance dashboard flags a 400% jump in a line item nobody touched. Run your own numbers on the calculator if your read-to-write ratio is nothing like four to one, because with a flat 5.00x it will not change the multiplier, only the magnitude.
The cheapest model changes hands twice in 33 hours
We swept every model on OpenRouter carrying a non-zero discount field and then read the banner on each hit. Two of them, and only two, are vendor-announced promotions with a published end date rather than reseller margin against list. Both expire this week, 33 hours apart, and they happen to sit next to each other at the bottom of the price table.
| From | Cheapest in the basket | Cost | What just lapsed |
|---|---|---|---|
| Now, through Sep 8 06:59 UTC | Mercury 2.5, discounted | $0.0775 | Both discounts live |
| Sep 8, 07:00 UTC | GLM-5.3-Flash, discounted | $0.1375 | Mercury reverts to $0.20 / $0.75 |
| Sep 9, 16:00 UTC | Qwen3.8-Flash | $0.2085 | GLM reverts to $0.15 / $0.50 |
Three different models hold the cheapest slot inside two days and not one of them changes. No retraining, no new checkpoint, no price cut, no price rise that anyone decided this week. Two clocks run out and the leaderboard reshuffles itself. If you have a router that picks a model on price, it will make three different decisions between now and Wednesday evening for reasons that have nothing to do with the models. Whatever Qwen3.8-Flash inherits on Wednesday it inherits by standing still, which is also how it lost the comparison to GLM-5.3-Flash when both landed in August.
The speed number, and the speed number
Inception sells Mercury on throughput, so the throughput claim deserves the same scrutiny as the price. The pitch, republished verbatim by OpenRouter, is 1,107 tokens per second "on standard GPUs". That is the entire hardware disclosure. No GPU model, no batch size, no concurrency, no reasoning level, no statement of whether it is a single stream or an aggregate. Mercury 2's equivalent claim at least named the silicon.
| Model | Vendor claim, tok/s | OpenRouter P50 | AA measured |
|---|---|---|---|
| Mercury 2.5 | 1,107, standard GPUs | 471 | no page |
| Mercury 2 | 1,009, NVIDIA Blackwell | 233 | 709.9 |
| Gemini 3.5 Flash-Lite | not published | - | 338.3 |
| Gemini 3.8 Flash, high | not published | - | 280.8 |
| GPT-5.4 nano, xhigh | not published | - | 178.3 |
| DeepSeek V4 Flash, max effort | not published | - | 131.8 |
| GPT-5.6 Luna, max | not published | - | 127.8 |
| Claude Haiku 4.5, non-reasoning | not published | - | 80.2 |
| GLM-5.3-Flash | not published | - | 51.5 |
Do not read across those columns. A vendor figure from a lab, a median across live production traffic, and a controlled independent harness are three different questions, and the honest comparison is column against itself. Within OpenRouter's own numbers, Mercury 2.5 serves 471 tokens a second against Mercury 2's 233. That is 2.02x, measured the same way on the same platform on the same day, and it is the only speedup claim in this post we would defend. The vendor claim sits at 2.35x the served median, so we would treat 1,107 as a ceiling on good hardware and 471 as what shows up.
Artificial Analysis has no page for Mercury 2.5 at all. It has one for Mercury 2, currently reading 709.9 tokens a second and an Intelligence Index of 15, which is second of 177 on speed and near the bottom on quality. We should flag a correction on ourselves here: our Mercury 2 post in May cited 788 tokens a second from Artificial Analysis and that page now reads 709.9. We cannot reconstruct the original capture, so we have updated our catalogue to the current figure and dated it. Measured speed is a moving number in a way a rate card is not, which is reason enough to date every one you publish.
Twice as fast at the same price is twice as expensive per hour
Here is the bit that catches people. Mercury 2 and Mercury 2.5 both charge $0.75 for a million output tokens at list. Mercury 2.5 emits about twice as many of them per second. Tokens are the billing unit, so an hour of continuous generation costs about twice as much on the faster model.
| Model | Tok/s served | Tokens per hour | Cost per hour |
|---|---|---|---|
| Mercury 2.5, discounted | 471 | 1,695,600 | $0.25 |
| Mercury 2 | 233 | 838,800 | $0.63 |
| Mercury 2.5, list | 471 | 1,695,600 | $1.27 |
None of this makes speed worthless. It makes speed worthless as a way to spend less on tokens. A per-token meter is indifferent to how quickly you hit it: the only thing 2x throughput buys you is the same bill arriving in half the time. Speed converts to money when finishing early lets you stop paying for something measured in seconds instead of tokens, which is exactly the workload Inception names in its own pitch. A voice pipeline where latency compounds, a search agent holding a container open, a coding subagent with a person waiting on it. If your job is a nightly batch that nobody is watching, you are paying a speed premium into thin air, and there is no batch tier here to opt out into.
The same logic runs through the reasoning levels. OpenRouter exposes four settings and there is no surcharge attached to any of them. Inception's FAQ is explicit that reasoning tokens bill at the ordinary output rate, and its API returns them as a subset of the completion count rather than a separate line. So turning effort up does not change your price per token by a cent. It changes how many tokens you buy, which is the same trick from the other direction.
Nobody has benchmarked this model
We went looking for a quality figure and came back with nothing. Inception has published no benchmark for Mercury 2.5 and no launch post to put one in. Artificial Analysis returns a 404 on every slug we tried. OpenRouter's model API carries a benchmarks object for Mercury 2, with Design Arena elo across eight categories and an Artificial Analysis coding index of 31.1, and carries nothing at all for 2.5. That absence is the cleanest evidence available that no independent measurement exists yet.
What exists is one sentence: "a 10+ point jump in intelligence over Mercury 2", with no index named. If the unnamed index happens to be the Artificial Analysis one, Mercury 2 sits at 15 and the claim would put 2.5 somewhere around 25, which would still be under Gemini 3.5 Flash-Lite's 28 and a long way under GPT-5.6 Luna's 43. That is our arithmetic on a guess about which index was meant, and we would not want it quoted as a score. It is worth doing only because it is the sole way to sanity-check the company's other claim, that Mercury 2.5 is "comparable quality to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5".
Be careful with the numbers circulating on aggregator sites this week. Several are quoting AIME 2025 at 91.1 and GPQA Diamond at 73.6 for Mercury 2.5. Those are Mercury 2 figures, and we could not find them on any Inception page either, so we would not attach them to the older model with confidence and we would certainly not attach them to the newer one. At $0.04 a million the absence of a score is an acceptable trade for most people running an evaluation of their own. At $0.20 it is a real question, because the models it is now priced against all publish theirs.
Every dated discount currently running
We built this ledger in August and it needs a new row roughly every fortnight. These are promotions the vendor itself announced with an end date attached, not marketplace resellers discounting against list.
| Model | Now | After | Published expiry |
|---|---|---|---|
| Inception Mercury 2.5 | $0.04 / $0.15 | $0.20 / $0.75 | Sep 8, 2026, 07:00 UTC |
| Z.AI GLM-5.3-Flash | $0.075 / $0.25 | $0.15 / $0.50 | Sep 9, 2026, 16:00 UTC |
| Google Gemini 3.8 Flash | $0.75 / $3.75 | $1.50 / $7.50 | Dec 31, 2026 |
| Google Gemini 3.7 Flash | $0.75 / $3.75 | $1.50 / $7.50 | Dec 31, 2026 |
| Google Gemini 3.6 Flash | $0.75 / $3.75 | $1.50 / $7.50 | Dec 31, 2026 |
| OpenAI GPT-5.6 Sol | $4.00 / $20.00 | not published | at least through Nov 21, 2026 |
Four different conventions in six rows. Inception gives you an hour but only through its distributor. Z.AI gives you an hour on its own docs, in Singapore time, and OpenRouter independently converts it to UTC. Google gives you a calendar date and the exact number that follows, which is why the January 1 doubling is the easiest of these to plan around despite being the furthest away. OpenAI gives you a date and withholds the price, which is the worst of the four, because a date without a number tells you when to be worried and not how much.
Anthropic sits outside the table with the outcome nobody prices in. Claude Sonnet 5 launched on introductory rates of $2 and $10 that were meant to lapse to $3 and $15 on September 1. Anthropic cancelled the increase and made the promotional number the standard one. A published expiry is a statement of intent, not a contract, and it can resolve downward. We would not bet a migration on that happening here, though, because Inception has already published the post-promotion price on its own website and has been showing it there all week.
Whether this is worth doing anything about
Not much, honestly, and that is the useful conclusion rather than a deflection. A 5.00x rise on a base of $0.0775 per basket is a rise to $0.3875 per basket. If Mercury 2.5 is a serious line in your budget you are running enormous volume, and if you are running enormous volume on a preview model with no published benchmarks and no vendor API you have a larger problem than Tuesday.
- If you were evaluating it, finish the evaluation today. The cost of an eval run goes up 5x tomorrow morning and the model does not change, so whatever you were going to learn is a fifth of the price for the next few hours.
- If you shipped on it, check whether your router is picking on price. Three different models take the cheapest slot in our basket between now and Wednesday evening, and a naive price-first router will chase all three.
- If you are on Mercury 2, move. Same output price, 20% off input, double the context, roughly double the served throughput. That was true last week and it stays true after the discount lapses.
- If you are shopping on price alone in this tier, Qwen3.8-Flash at $0.113 and $0.382 is the one still standing on Wednesday, and Alibaba's page says explicitly that those figures exclude promotions, which is a sentence we wish more vendors printed.
The wider point is the one that keeps recurring on this site. The sticker price of a model is increasingly a schedule rather than a number, and the schedule lives somewhere different for every vendor: a footnote, a banner on a marketplace, a struck-through figure, a paragraph in a docs page, a timezone you have to convert. Our catalogue stores the price in effect and the price that replaces it, with the date, because a calculator holding only one of those goes silently wrong on a morning nobody has in their diary. This particular morning is tomorrow.
Questions we got asked while writing this
Is the September 8 date reliable?
It is as reliable as this gets. The banner on OpenRouter reads "Limited-time 80% discount via Inception through September 8, 2026 at 07:00 UTC", and OpenRouter is the merchant of record. Inception's own post on X is indexed with the phrase "80% off pricing through 9/7", which is the same instant written in Pacific wall-clock, since midnight PDT on September 7 is 07:00 UTC on September 8. We could not open the post itself, so we are relying on the distributor's timestamp, which is the precise one anyway.
Will the cache discount soften the increase?
No. The cache read rises by the same 5.00x, from $0.004 to $0.02. It stays a 90% discount against input on both sides of the date, so a team caching heavily pays 5x more than it does today, exactly like a team caching nothing. There is no cache write price published at all, and no batch tier to fall back on.
Can I buy Mercury 2.5 anywhere other than OpenRouter?
Not today. Inception's own API returns exactly one model and it is Mercury 2. If a second channel opens after the promo lapses, the list price is already published, so at least you know what to expect. Mercury Coder, by the way, is gone from Inception's site, its docs, its API and OpenRouter's whole catalogue. Third-party listings persist. We would not read those as evidence you can still buy it.
Is a diffusion model priced differently from a normal one?
Not in any way you can see on the bill. Mercury generates and refines many tokens in parallel rather than one after another, but it still meters input tokens, output tokens and cached reads, exactly like an autoregressive model. The architecture shows up in throughput, not in the shape of the invoice, which is why the per-hour arithmetic above works out the way it does.
Where these numbers came from, and where they run out
- OpenRouter: Mercury 2.5 Preview - The promotional banner and its 07:00 UTC timestamp, the struck-through $0.20, $0.75 and $0.02, the served throughput of 471 tokens a second at the median and 0.77s latency, and Inception's model description including the 1,107 tok/s claim and the comparison to Luna, Flash-Lite and Haiku. The endpoints API returns an explicit
discount: 0.8field; the same endpoint for Mercury 2 returnsdiscount: 0 - Inception: models and the docs pricing table - The $0.20, $0.02 and $0.75 list card for Mercury 2.5 with no discount and no date; Mercury 2 at $0.25, $0.025 and $0.75 with a 128K window and 50,000 max output. The docs table does not contain a Mercury 2.5 row
- Inception: FAQ and Introducing Mercury 2 - That completion tokens including reasoning tokens bill at the output rate, that reasoning tokens are returned as a subset of the completion count, and the 1,009 tok/s on NVIDIA Blackwell figure for Mercury 2
- Z.AI: pricing - GLM-5.3-Flash at $0.075, $0.015 and $0.25 against a list of $0.15, $0.03 and $0.50, and the sentence "The promotion ends at 24:00 on September 9, 2026 (UTC+8, Singapore time)", which OpenRouter independently renders as 16:00 UTC on the same day
- Google: Gemini API pricing, OpenAI: pricing and Anthropic: pricing - Gemini 3.8, 3.7 and 3.6 Flash at $0.75 and $3.75 through December 31, 2026 and $1.50 and $7.50 from January 1; Gemini 3.1 Flash-Lite at $0.25 and $1.50 and 3.5 Flash-Lite at $0.30 and $2.50; GPT-5.6 Luna at $0.20 and $1.20 short context, GPT-5.4 nano at $0.20 and $1.25, and the note that Sol's promotional pricing runs at least through November 21, 2026; Claude Haiku 4.5 at $1.00 and $5.00, and Anthropic's statement that the scheduled Sonnet 5 increase will not occur
- DeepSeek: pricing and Alibaba Cloud: Qwen3.8-Flash - V4 Flash at $0.22 and $0.66 off-peak, $0.44 and $1.32 at peak, with peak defined as 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays; Qwen3.8-Flash at $0.113 and $0.382, on a page that states it shows original pricing excluding limited-time promotions. We deliberately did not blend DeepSeek's own rates with the cheaper reseller rates for the same open weights
- Artificial Analysis: Mercury 2 - 709.9 output tokens a second, second of 177 on speed, and an Intelligence Index of 15, plus the measured speeds for every other model in the speed table. Every Mercury 2.5 slug we tried returns a 404
- TokenCost: Mercury 2 pricing, the discount-expiry ledger and GLM-5.3-Flash against Qwen3.8-Flash - Our earlier work on this model, on dated price rises, and on the two budget models that inherit the top of the table this week
- What we could not establish. No benchmark score for Mercury 2.5 exists from any source, so the "10+ point jump" claim cannot be checked and does not name its index; our estimate of roughly 25 on the Artificial Analysis scale is arithmetic on a guess and should not be quoted as a measurement. The 1,107 tok/s figure comes with no hardware, batch size, concurrency or reasoning level attached, and we could not determine whether it describes a single stream or an aggregate. We could not open Inception's announcement on X, which returned HTTP 402 on every attempt, so the "through 9/7" wording is a search-indexed page title rather than something we read. The reasoning-effort levels differ between the two places they are documented, none/low/medium/high on OpenRouter against instant/low/medium/high in Inception's docs, and no Inception page documents Mercury 2.5's levels at all. Inception contradicts itself tenfold on its free tier, 100 million tokens in the docs and 10 million on the models page, so we left the free tier out of every calculation. OpenAI publishes an expiry date for GPT-5.6 Sol but not the rate that follows it, so that row has a date and no number. We could not retrieve Moonshot's per-model price tables, so Kimi is absent from the promo sweep rather than confirmed clear. OpenRouter's throughput figures are medians over live production traffic and will move with load, which is why we compared them only against each other. And our own May post cited 788 tokens a second for Mercury 2 from Artificial Analysis, whose page now reads 709.9; we could not reconstruct the original capture, so we have updated the catalogue and dated it rather than defend the older number.