Eight AI API discounts expire on a published date between September 7 and January 1. Half of them are video models nobody is tracking, the loudest one is the smallest, and a ninth that everybody lists has no end date at all.
Introductory pricing has quietly become the normal way to launch a model, which means a growing share of the rates in your spreadsheet are not prices so much as offers with a clock on them. We went looking for every one carrying a provider-published expiry date. Eight clear that bar, and they split exactly down the middle: four price tokens and four price video. There is also a ninth that gets listed everywhere, including in our own catalogue until this morning, which turns out not to have a deadline at all.

Photo by Declan Sun on Unsplash
What one unchanged stack does over the next four months
+$405 a month
$605 today, $1,010 in January, on two dates the providers have already printed. Nothing about the workload changes: same three models, same token counts, same code.
Monthly volumes: 300M input and 40M output on Gemini 3.7 Flash, 200M and 60M on GLM-5.3-Flash, 20M and 6M on GPT-5.6 Sol. Sol stays at its current $4/$20 throughout, for a reason we get into below. September 7, 17, 23 and 25 are missing from this ladder because those four are video models and this stack does not render any.
Everything with a date on it
The bar for this table was deliberately narrow: the provider has to have published the date itself. Not a launch-window guess, not a gateway showing a discount field, not a reporter's inference. Eight rows clear it. The ninth is in the table because it is on everyone else's version of this list, and it should not be.
| Date | Model | Today | After | Change |
|---|---|---|---|---|
| September 7, 2026 | Doubao-Seedance 2.0 mini | 40% of list, 480p and 720p | List, from 14:00 UTC+8 | 2.50x |
| September 7, 2026 | Doubao-Seedance 2.0 fast | 75% of list, 480p and 720p | List, from 14:00 UTC+8 | 1.33x |
| September 9, 2026 | GLM-5.3-Flash | $0.075 / $0.25 | $0.15 / $0.50 | 2.00x |
| September 10, 2026 | Solar Pro 4 | $0.03 / $0.12 | $0.30 / $1.20 | 10.00x |
| September 17, 2026 | Doubao-Seedance 2.5 | 72% of list, 1080p only | List, from 14:00 UTC+8 | 1.39x |
| September 23, 2026 | Wan 3.0 (Standard) | $0.14 / sec at 1080p | $0.20 / sec at 1080p | 1.43x |
| September 25, 2026 | Ling 3.0 Flash Fin (Vercel) | Free | $0.06 / $0.18 | from zero |
| January 1, 2027 | Gemini 3.6 & 3.7 Flash | $0.75 / $3.75 | $1.50 / $7.50 | 2.00x |
| No published date | GPT-5.6 Sol | $4.00 / $20.00 | Not announced | unknown |
Per-token rates are input and output per million. The Seedance and Wan rows are video, priced per output second or per million tokens of a prompt that contains video, which is the first thing worth noticing: these are not all denominated in the same unit, so they cannot share a spreadsheet column even though they hit the same invoice.
The second thing is the September 9 and September 10 rows, which are consecutive days, and the quieter of the two is five times the step-up. GLM-5.3-Flash doubles and got a week of coverage; Upstage's Solar Pro 4 goes up tenfold the following evening and we could not find anyone writing about it at all.
Four of the eight are video, and that is not a coincidence
We went looking for dated promotions across Baidu, Tencent, Zhipu, ByteDance, Ant, Meituan, Xiaomi, StepFun and several others, and the result was lopsided enough to be the finding rather than a footnote. Excluding GLM-5.3-Flash, which we already had, there is not one currently-running dated promotion on a text model from any of them. Every dated Chinese promotion we could confirm is video generation.
ByteDance is running three at once on Volcengine, all ending at 14:00 Beijing time on a stated day. Seedance 2.0 mini bills 480p and 720p at 40% of list until September 7, which is the steepest discount in this entire piece and reverts to two and a half times what you pay today. Seedance 2.0 fast runs at 75% of list to the same deadline. Seedance 2.5 discounts only 1080p, to 72% of list, and runs ten days longer. Volcengine publishes these in yuan with a floor quoted per second, roughly 0.2 CNY/sec at 720p on the mini and 2.7 CNY/sec at 1080p on 2.5, and we are leaving them in yuan because converting them would mean picking an exchange rate and pretending it was part of the price.
Why the split falls this way is worth a guess, clearly labelled as one. Video inference is expensive enough per unit that a launch discount is a real marketing cost, so vendors bound it with a date and let it lapse. Text pricing has been falling for two years on its own, so a lab that wants to be cheap can simply be cheap and change the card again later. The practical consequence is not a guess: if you are tracking expiring rates and you only watch token pricing, you are watching the half of this calendar where less is happening.
September 9 got the coverage and moves 7.4% of the money
GLM-5.3-Flash launched on August 26 at $0.075 per million input tokens, which is half of Z.ai's own list price of $0.15, and the promotion behind that number ends at 24:00 UTC+8 on September 9. We wrote about it yesterday and so did most people, because it is a genuinely striking number attached to a genuinely good model. On the stack above it is worth $30 a month.
There is a wrinkle underneath it that the coverage mostly skipped. Pull OpenRouter's endpoint list for the model and fifteen sellers come back, of which exactly three carry the promotional rate: Z.AI itself, Novita and GMICloud. Eleven are already charging the full $0.15 and $0.50, and Modal has an unrelated discount of its own that lands within a hundredth of a cent of list. So the promotion is not really a market price that expires on September 9. It is a price three vendors are running while twelve of their competitors quietly charge what the card says, and the deadline merely ends the disagreement.
The two nobody has written up are the steepest on the page
Upstage is selling Solar Pro 4 at $0.03 per million input tokens and $0.12 output, against a list card of $0.30 and $1.20 on its own pricing page. That is 90% off, it applies on Upstage Console and on Upstage's first-party OpenRouter endpoints, and it ends at 23:59 UTC on September 10. On the same 300M input and 40M output leg we have been using, that is $13.80 becoming $138.00. Ten times, to the cent, and the steepest step-up we found anywhere. One inconsistency to note if you are buying through the gateway rather than the console: the sentence on Upstage's page that mentions OpenRouter gives the date without a timezone, and only the console sentence carries the 23:59 UTC.
The other one is structurally the most interesting thing in this whole calendar. Vercel is running Ling 3.0 Flash Fin free on its AI Gateway through September 25, after which it bills at $0.06 and $0.18. What makes it worth a paragraph is that Vercel shipped two model IDs rather than one. Point at inclusionai/ling-3.0-flash-fin and your requests quietly start costing money on September 26. Point at inclusionai/ling-3.0-flash-fin-free and the same requests start returning an error instead. Vercel is letting you choose in advance whether the end of a free tier is a bill or an outage, which we have not seen another provider offer and which is a genuinely good idea. Going from free to any positive number is the one row on this page with no meaningful ratio, and also the only one where the failure mode is a surprise invoice rather than a larger one.
January 1 got almost none and moves 92.6%
Google has had the January 1 step-up on its own pricing page since Gemini 3.7 Flash went GA on August 13. It is not hidden and it is not new. It is written into the model rows as dated text: $0.75 through December 31, then $1.50 from January 1, with output moving $3.75 to $7.50 the same day. Gemini 3.6 Flash carries an identical pair of rows, and Google's Vertex page names CodeMender in the same sentence, so three products move together.
Of the $405 a month our example stack picks up on published dates, $375 of it is this one line. That is 92.6%, against 7.4% for the expiry that dominated the week's coverage. The asymmetry is not a quirk of the volumes we picked either. It falls out of two things that tend to travel together: cheap tiers attract the highest token counts, and Google chose to double rather than trim.
One framing correction while we are here, including to how we have described this ourselves. Calling the current rate a 50% cut is wrong. Neither Flash model has ever been sold at $1.50 and $7.50. They launched at $0.75 and $3.75 and that is the only rate either has ever charged, so nothing is being reversed on January 1. A price that has never existed is scheduled to start existing, which is a different and slightly stranger thing than a discount lapsing.
You cannot cache your way out of a doubling
The instinct when a rate doubles is to reach for the discount levers, and Gemini has good ones. It is worth doing the arithmetic before you plan around them, because they all move together. Here is the same 300M input and 40M output leg on all four billing modes Google publishes, on both sides of January 1.
| Billing mode | Through Dec 31 | From Jan 1 | Ratio |
|---|---|---|---|
| Standard | $375.00 | $750.00 | 2.00x |
| Cached input, 90% off | $172.50 | $345.00 | 2.00x |
| Batch, 50% off | $187.50 | $375.00 | 2.00x |
| Vertex Priority, 1.8x | $675.00 | $1,350.00 | 2.00x |
Four modes, one ratio. Caching is a percentage off the input rate, Batch is a percentage off both, and Priority is a multiplier on both, so doubling the base doubles every line derived from it. Cache storage doubles too, $0.50 to $1.00 per million tokens per hour, which on a million tokens held for a day is $12 becoming $24. The levers are worth pulling and several of them are worth pulling today. None of them is a response to January 1.
What does change the ratio is running the work somewhere else, and the reason this cliff matters is that the somewhere else is not marginal. Priced on that same leg:
| Model | 300M in, 40M out | Against Gemini |
|---|---|---|
| Gemini 3.7 Flash, from January 1 | $750.00 | the baseline |
| GLM-5.3-Flash, list card | $65.00 | 11.54x cheaper |
| Qwen3.8-Flash | $66.80 | $66.80, a nose behind |
| DeepSeek V4 Flash, off-peak | $92.40 | under an eighth |
Note what happened to the September 9 expiry in that table. Even after GLM-5.3-Flash goes back to full list price, it is still 11.5 times cheaper than Gemini Flash will be five weeks later. The cliff everyone is bracing for lands the cheaper model in a position it was already going to win from. Whether an 11.5x gap is worth taking depends entirely on whether the two models do your job equally well, which is a question no price table can answer and which we would not assume either way.
“At least through” is not an expiry date
OpenAI cut GPT-5.6 Sol from $5/$30 to $4/$20 on August 21, and the sentence attached to it reads that the promotional pricing is available at least through November 21, 2026. That phrasing has been widely reproduced as a November 21 expiry, ours included. It is the opposite. A floor is not a ceiling. November 21 is the earliest date on which OpenAI has reserved the right to move the price, and OpenAI has published no rate for November 22 and no end date whatsoever.
We had this wrong in a way worth naming precisely, because it had teeth. Our catalogue stores scheduled changes as a structured field with an end date and a replacement price, which is the right shape for Gemini and for GLM-5.3-Flash. We filled the same field in for Sol with November 21 and a reversion to $5/$30, on nothing more than the absence of a published alternative. That would have made the calculator raise a customer's Sol estimate by 40% on November 22 on OpenAI's behalf, for a change OpenAI has never announced. It is removed as of this post. The note explaining the November 21 floor stays, because the floor is real and it is genuinely useful to know; the invented cliff does not.
That $80 a month gap between $1,010 and $1,090 in the opening arithmetic is the whole size of the assumption. It is 16.5% of what a naive reading of the calendar would tell you to budget for, and it rests on nothing. There is a second Sol discount worth flagging for the same reason: OpenRouter is currently running its own 50% cut on the OpenAI routes, taking Sol to $2/$10, half of OpenAI's already-promotional rate. It has been live since roughly August 17 and it has no published end date either. If your bill depends on it, that dependency is on somebody's pricing decision rather than on a calendar.
September 23 is billed in seconds, and two platforms disagree
Alibaba released Wan 3.0 on August 23 and it is the one row here that does not touch a token. Video is billed per output second, on a Standard tier at $0.05, $0.10 and $0.20 per second for 480p, 720p and 1080p, and a Prime tier at $0.068, $0.14 and $0.28. A thirty-second 1080p clip is therefore $6.00 at Standard list, and the launch offer that runs to September 23 takes 30% off to make it $4.20.
Where you buy it decides how much of that you get. Alibaba's own announcement scopes the 30% to Standard tier on Alibaba Cloud Model Studio and Qwen Cloud. Buy the same model through OpenRouter and the discount is 15%, which is exactly half, and puts the same clip at $5.10. Prime is not discounted anywhere we could find. We would guess the 15% is a partial pass-through of the 30%, and we are flagging that as a guess: nobody has published why the two numbers differ, and a plausible mechanism is not a source. For a team rendering two hundred thirty-second 1080p clips a month, the September 23 date is $840 becoming $1,200.
The one that got called off
A fifth date belongs on this page precisely because it is not happening. Claude Sonnet 5 launched with introductory pricing of $2/$10 running through August 31, 2026, and a scheduled increase to $3/$15 on September 1. On August 10 Anthropic said the introductory rate is now the standard price and the September 1 increase will not occur. Anthropic's pricing documentation carries the same note, and Sonnet 4.6 and 4.5 sit above it on that page still charging the $3/$15 that Sonnet 5 was supposed to move to.
We covered that cancellation at the time and removed the schedule from the catalogue so the calculator would not raise anyone's Sonnet 5 estimate on September 1. It is slightly uncomfortable to notice that we applied exactly that discipline to Anthropic in August and failed to apply it to OpenAI in the same catalogue three entries down. The lesson generalises better than the individual dates do: a published cliff is a stated intention, competitive conditions can retire it, and the honest way to store one is with the source attached so you can tell the difference between a date a provider printed and a date somebody inferred.
Two dates are worth your afternoon. Six are not.
If you take one thing from the arithmetic, make it the ranking rather than any individual figure. September 9 is close, loud, and worth 7.4% of our example stack. January 1 is distant, unmentioned, and worth 92.6%. September 10 is a tenfold step-up on a model we found no coverage of whatsoever. The urgency of a deadline, the volume of writing about it, and the size of it turn out to be close to unrelated, and the ordering has been wrong in the same direction all month.
The practical version is short. Find out which of your lines are introductory rates rather than prices, which is not obvious from any dashboard we know of and generally requires reading the provider's own page. Price your logged token counts on the replacement card rather than the current one, because that is the number your January invoice uses. Then check whether the levers you are counting on survive the change, which for Gemini they do not in any useful sense. And store the date next to the price instead of overwriting one with the other; if that sounds fussy, it is the difference between the $1,010 in the opening card and the $1,090 we would have told you a week ago.
Everything above is in the pricing table with the current rate, the replacement rate and the provider page we read them from stored as separate fields, and with one fewer invented deadline in it than there was yesterday. If you want to run your own volumes against both cards, the calculator takes them directly.
Eleven pages, three APIs, and the five vendors we did not reach
- Google: Gemini API pricing - The January 1, 2027 step-up, written as dated rows on both Gemini 3.7 Flash and Gemini 3.6 Flash: $0.75 input through December 31 then $1.50, $3.75 output then $7.50, cached input $0.075 then $0.15, cache storage $0.50 then $1.00 per million tokens per hour. Google's Vertex AI pricing page states it once as prose, uses the word introductory, and names CodeMender alongside both models. Batch moves $0.375/$1.875 to $0.75/$3.75; Vertex non-global endpoints carry a 10% premium and double as well
- Z.ai: model pricing - The 50% GLM-5.3-Flash discount and its deadline of 24:00 on September 9, 2026, Singapore time, in a single sentence, with the list card printed as strikethrough beside the promotional one. Promotional $0.075/$0.015/$0.25 against list $0.15/$0.03/$0.50
- OpenRouter: GLM-5.3-Flash endpoints - Pulled August 28, 2026, no key required. Fifteen sellers, of which Z.AI, Novita and GMICloud return a discount field of 0.5 and quote the already-discounted rate. Eleven return 0 and charge list. Modal returns 0.6667, which resolves to $0.149985 and $0.49995, within a hundredth of a cent of list and unrelated to Z.ai's promotion. The same API is the source of OpenRouter's own 50% cut on the OpenAI Sol routes
- OpenAI: 20% price reduction for GPT-5.6 Sol - Posted August 21, 2026, says the cut applies starting today, $5 to $4 on input and $30 to $20 on output. The critical phrase for this post is that the promotional pricing is available at least through November 21, 2026, which is a floor and not an expiry. OpenAI's pricing page confirms the current $4.00/$0.40/$20.00 short-context card and $8.00/$0.80/$30.00 long-context card
- The Decoder: Alibaba's Wan 3.0 - The per-output-second card at both tiers: Standard $0.05, $0.10 and $0.20 for 480p, 720p and 1080p, Prime $0.068, $0.14 and $0.28. Alibaba's own account scopes the 30% launch offer to Standard tier on Alibaba Cloud Model Studio and Qwen Cloud from August 23 to September 23; TechNode reports the same discount as running from August 24 on selected platforms and carries no per-second figures of its own. OpenRouter's single Wan 3.0 endpoint returns a discount of 0.15
- Anthropic: pricing - The note stating that Sonnet 5's $2/$10 introductory rate is now standard and the September 1, 2026 increase to $3/$15 will not occur. Anthropic's Sonnet 5 announcement carries the update dated August 10, 2026
- Volcengine: model pricing - All three Doubao-Seedance promotions, each written as a dated sentence with a 14:00 UTC+8 start and end. Seedance 2.0 mini and 2.0 fast run August 7 to September 7, at 40% and 75% of list on 480p and 720p; Seedance 2.5 runs August 14 to September 17 at 72% of list on 1080p only, with 480p and 720p explicitly excluded. List rates are in yuan per million tokens, quoted separately for prompts with and without video, with per-second floors of roughly 0.2 CNY at 720p on the mini and 2.7 CNY at 1080p on 2.5
- Upstage: Solar Pro 4 - The 90% launch promotion and its September 10 deadline. The page states it twice, once as 90% off on Upstage Console through September 10 at 23:59 UTC, and once as 90% off on Console and OpenRouter through September 10 with no timezone given. List prices of $0.30 input, $1.20 output and $0.06 cached input come from Upstage's pricing page, and OpenRouter's endpoints API independently returns a discount field of 0.9 on both of Upstage's own endpoints
- Vercel: Ling 3.0 Flash Fin free on AI Gateway - Changelog dated August 27, 2026. Free through September 25, then the standard rate of $0.06 input, $0.18 output and $0.01 cache read. Source of the two-model-ID mechanism: the plain ID begins billing when the offer ends, and the -free suffixed ID returns an error instead. No timezone is published for the September 25 date. Vercel's discounted-model list returns exactly two entries today, this one and Gemini 3.7 Flash at 50% off through December 31, and we have deliberately not counted that second one separately because it is Google's own January 1 step-up restated at gateway level rather than an independent promotion
- TokenCost: two models called Flash - Our August 27 post, source of the GLM-5.3-Flash and Qwen3.8-Flash rate cards used in the alternatives table, and of the observation that the September 9 date is worth more than the rate-card gap it hides
- What this list does and does not cover. We verified OpenAI, Google, Anthropic, Z.ai, Alibaba, ByteDance, Upstage and Vercel against first-party pages, and separately confirmed that Cohere, AWS Bedrock, Perplexity, Together, Fireworks, Groq, Baseten, Cerebras, Baidu, Tencent, Xiaomi, StepFun and Meituan have nothing dated currently running. We did not reach primary pricing docs for DeepSeek, Moonshot, MiniMax, Mistral or xAI, so read this as thorough rather than exhaustive. Three near-misses are worth naming because they look like they belong and do not: Ant Group's Ling-3.0-flash 75% discount says only that list prices resume gradually from September, with no day and no timezone, and the September 1 date circulating in secondary coverage is not on Ant's page; Meituan's LongCat-2.0 launch discount and Ant's 90% cut across the 2.6 series are both real and both undated; and a mechanical sweep of every model on OpenRouter turned up 36 endpoints carrying a non-zero discount field, none of which publishes an expiry. A discount field proves a price gap, not a deadline, which is also true of Vercel, where the rates sit in the API and the dates exist only in changelog prose. Finally, the reason OpenRouter passes through 15% of Alibaba's 30% is documented nowhere we looked, so the partial pass-through in the Wan 3.0 section is our inference, and the explanation offered above for why dated promotions cluster in video is a guess we have labelled as one