Qwen3.7 Flash lists at $0.03 a million. One token past 32K, the identical request bills 3.3x more.
The sticker is real and it is the lowest one attached to any 1M-context model you can call this morning. It is also one of three, because Alibaba prices this model in steps: $0.03 and $0.13 per million up to 32K of input, $0.10 and $0.40 to 256K, $0.20 and $0.80 beyond that. The steps are not marginal bands. Send 32,001 input tokens instead of 32,000 and the whole request reprices, which makes that one token cost more than the 32,000 before it. Blend the middle tier at a normal 3 to 1 ratio and you land on $0.175 per million, the same figure DeepSeek V4 Flash blends to on a card with no steps in it. Past 256K, DeepSeek is simply cheaper on both sides. And the thing that would settle the argument does not exist: Alibaba published no benchmarks for this model, Artificial Analysis has not measured it, and the weights are closed, so nobody outside Hangzhou can tell you what the $0.03 buys.

Photo by Mike Hindle on Unsplash
What you are actually buying
A very cheap short-prompt model with a 1M context window it charges a lot to use. Those two facts sit awkwardly together, because the workloads that want a million tokens of context are exactly the ones that never see the advertised rate. If your prompts are short, this is the cheapest thing on the market by a wide margin and you should try it. If they are not, the honest comparison is against DeepSeek and Mistral, not against the $0.03.
What the label says
$0.03 / $0.13
and it holds only while your input stays under 32K
What one token past 32K says
$0.10 / $0.40
applied to the whole request, not the overflow
Three prices, and only one of them is on the label
Alibaba's own pricing tables are headed Input<=32k, 32k<Input<=256k and 256k<Input<=1m. Read that literally, because it matters: the tier is chosen by input length, and the output price sits inside the input tier. Your completion length never moves you between bands, and a short reply to a long prompt is billed at long-prompt output rates.
| Input length | Input | Output | Cached read | 3:1 blend |
|---|---|---|---|---|
| Up to 32K | $0.03 | $0.13 | $0.003 | $0.055 |
| 32K to 256K | $0.10 | $0.40 | $0.010 | $0.175 |
| 256K to 1M | $0.20 | $0.80 | $0.020 | $0.350 |
All figures are USD per 1M tokens on the international endpoint. Cached read shown is the explicit cache rate at 10% of input; the implicit cache reads at 20%. Top to bottom, input multiplies by 6.67 and output by 6.15, which is a wider internal spread than most people carry between two entirely different models.
The expensive token is number 32,001
Here is the same request either side of the first boundary, with a 1,000-token completion in both cases.
| Request | Rate applied | Cost |
|---|---|---|
| 32,000 in / 1,000 out | $0.03 / $0.13 | $0.001090 |
| 32,001 in / 1,000 out | $0.10 / $0.40 | $0.003600 |
| 256,000 in / 2,000 out | $0.10 / $0.40 | $0.02640 |
| 256,001 in / 2,000 out | $0.20 / $0.80 | $0.05280 |
The first jump is 3.30x and the second is exactly 2x. My favourite way to state the first one: a 32,001-token request costs what a 115,670-token request would have cost if the bottom tier had held. You are paying first-tier prices for a prompt more than three and a half times longer than the one you actually sent, and then not getting it.
Cliffs like this are not new. GPT-5.6 doubles input and adds 50% to output past 272K, which we wrote about when GPT-5.5 introduced it. What makes Qwen's version bite harder is where the first one sits. 272K is a deliberate choice you make when you decide to stuff a repository into a prompt. 32K is roughly a medium PDF, twenty turns of chat history, or one moderately chatty agent loop, and most teams cross it without ever deciding to.
Where the middle tier actually lands
Blend each tier at three parts input to one part output and put it next to the rest of the budget shelf. The middle row is the one worth staring at.
| Model | Input | Output | 3:1 blend | AA Intelligence |
|---|---|---|---|---|
| Qwen3.7 Flash, under 32K | $0.03 | $0.13 | $0.055 | not measured |
| Qwen3.7 Flash, 32K to 256K | $0.10 | $0.40 | $0.175 | not measured |
| DeepSeek V4 Flash 0731 | $0.14 | $0.28 | $0.175 | 49.93 |
| Mistral Small 4 | $0.15 | $0.60 | $0.263 | 19.56 |
| Qwen3.7 Flash, 256K to 1M | $0.20 | $0.80 | $0.350 | not measured |
| GPT-5.6 Luna | $0.20 | $1.20 | $0.450 | 51.24 at max |
| Qwen3.7 Plus | $0.40 | $1.60 | $0.700 | 38.98 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $0.850 | 36.48 |
Qwen3.7 Flash's middle tier and DeepSeek V4 Flash blend to the same $0.175. That is a coincidence of arithmetic rather than a conspiracy, but it is a useful one, because it strips price out of the decision entirely and leaves you comparing a measured 49.93 against a blank. DeepSeek also holds that rate flat across its whole 1M window, so the moment your prompts drift above 256K the tie breaks and does not come back. One asterisk on that flatness: DeepSeek has published a peak-hour surcharge that would double every line during Beijing working hours, with no start date attached. It is not being charged today, but it is written down.
Three notes on that table. GPT-5.6 Luna posts six different scores at six reasoning efforts for one price, and 51.24 is the top of that range rather than what you get by default. Mistral Small 4 sits well below everything else here on measured intelligence, so read its cheap blend accordingly. And Qwen3.7 Plus is itself tiered, doubling to $1.20 and $4.80 above 256K, so the $0.40 row is its short-prompt price in the same way the $0.03 row is Flash's. OpenRouter currently passes through a promotional $0.32 and $1.28 for Plus, which is a discount on that list rather than a different card. Prices came from each provider's own documentation and the scores from Artificial Analysis; you can see the full board on our pricing page.
Three workloads, one million requests each
Blended rates hide the cliff, so here is the same month of traffic under three different prompt shapes. Each row is a million requests, uncached, at list.
| Workload | Qwen3.7 Flash | DeepSeek V4 Flash | Mistral Small 4 | GPT-5.6 Luna |
|---|---|---|---|---|
| Ticket triage8K in / 500 out | $305 | $1,260 | $1,500 | $2,200 |
| RAG lookup60K in / 1.5K out | $6,600 | $8,820 | $9,900 | $13,800 |
| Transcript digest120K in / 2K out | $12,800 | $17,360 | $19,200 | $26,400 |
| Repo-wide agent pass400K in / 4K out | $83,200 | $57,120 | n/a256K context | $84,800 |
On short prompts it is a rout. DeepSeek costs four times as much for that month and Luna over seven, which on this volume is $955 a month against $1,895 of difference you simply keep. Nothing else on the shelf is close, and if your product is a classifier, a router or a summariser of small chunks, stop reading and go run a bake-off. The middle two rows tell a duller story: once the cliff has fired, the saving against DeepSeek settles at roughly a quarter and stays there whether you send 60K or 120K. By the bottom row it has inverted. The cheapest-listed model on the market costs 46% more than DeepSeek for the same repo pass, and lands within 2% of GPT-5.6 Luna, which does have a measured score of 51.24. Mistral Small 4 is absent from that row rather than winning it, because a 256K context window cannot accept a 400K prompt at any price. Run your own token mix through the calculator rather than trusting any of these three shapes to match yours.
The usual escape hatches are half shut
When a rate card has a cliff in it, caching is the normal way out. Qwen supports both flavours: implicit cache reads bill at 20% of input, explicit reads at 10%, explicit writes at 1.25x. Those are percentages of whichever tier you are in, so caching reduces your bill but never moves you down a step. A cached 300K-token prefix still bills at 10% of $0.20, not 10% of $0.03.
And a 90% explicit discount is just the going rate now, which we covered when surveying caching across every major API. DeepSeek gives 98%, and the arithmetic is unkind: DeepSeek's cache hits cost $0.0028 per million against Qwen's best cached rate of $0.0030. On the one line item where you would expect the cheap model to be untouchable, it is second.
Batch is worse. Alibaba documents a 50% discount for batch file jobs, then marks batch inference unsupported everywhere except Beijing. If you call the Singapore, Frankfurt or Tokyo endpoint you have no batch tier at all, while Gemini, OpenAI, Anthropic and Mistral all hand you 50% off for accepting a queue. The docs are also internally inconsistent here, repeating the Beijing price table under regions they have just declared unsupported, so confirm against your own console before planning around it. Fine-tuning is unsupported in every region.
Nobody has measured this model
I went looking for a number to put against these prices and came back with nothing. There is no Qwen blog post for Qwen3.7 Flash. There is no technical report. The Model Studio card is one paragraph of marketing copy that mentions multimodal understanding, spatial intelligence and a smoother vibe coding experience, with no evaluation table anywhere on the page. Alibaba has published zero benchmark scores.
Artificial Analysis has not filled the gap either. Qwen3.7 Flash does not appear in its model list, its leaderboard payload, or on its Alibaba provider page, which carries exactly two Qwen3.7 entries: Max and Plus. So there is no Intelligence Index, no coding index, no measured output speed, and no cost-per-task figure. Do not let anyone hand you a score for qwen3-5-omni-flash and call it this model.
Normally the open-weights community fills that vacuum within a week. Not here. The Qwen3.7 generation is closed: a Hugging Face API query for Qwen3.7 under the Qwen organisation comes back empty, and OpenRouter carries a single first-party endpoint. That last detail is at least a small mercy for price comparison: with no resellers there are no quantised copies undercutting the list price, which is the trap that makes OpenRouter's headline figures for DeepSeek and GLM misleading. We covered the same silence when Qwen3.8 Max shipped without benchmarks two weeks ago. Twice in a month is starting to look like a policy rather than an oversight.
The endpoint you call changes the price by up to a quarter
Alibaba publishes this card in CNY across four regions, and they are not the same card. Beijing, Frankfurt and Tokyo pay 0.2 and 0.8 CNY per million in the bottom tier. Singapore, the endpoint most international traffic lands on, pays 0.225 and 0.974. The USD prices everyone quotes convert from the Singapore table at about 7.5 CNY to the dollar, so the $0.03 headline is already the marked-up one.
| Tier | Beijing, in USD | Singapore, in USD | Uplift |
|---|---|---|---|
| Up to 32K | $0.0267 / $0.1067 | $0.03 / $0.13 | 13% in, 22% out |
| 32K to 256K | $0.0800 / $0.3200 | $0.10 / $0.40 | 25% on both |
| 256K to 1M | $0.1600 / $0.6400 | $0.20 / $0.80 | 25% on both |
USD equivalents are converted at the same 7.5 rate the official USD card implies, so treat them as comparison figures rather than something you can pay. Frankfurt and Tokyo sit on the Beijing numbers, which is odd and worth checking on your own invoice, because it means the cheapest legal way for a European team to call this model may be the Frankfurt endpoint rather than the international default. Rate limits differ too: 5M tokens a minute everywhere, but 30,000 requests a minute in Beijing against 15,000 in the other three.
The specs, minus the marketing
| Spec | Value |
|---|---|
| Context window | 1,000,000, of which 991,808 usable as input |
| Max output | 65,536 |
| Thinking mode | Same model, same price, 983,616 max input and a 262,144 reasoning budget |
| Modality | Text, image and video in, text out |
| Weights | Closed, one first-party endpoint |
| Availability | Snapshot dated July 15, listed publicly the week of July 27 |
One caveat on max output. Alibaba's Model Studio doc and OpenRouter both say 65,536, while the QwenCloud model page says 131K. That 131K happens to be the figure OpenRouter carries for Qwen3.7 Plus, which is the sort of thing that happens when a spec block is copied between pages, so I would plan against 65,536 and be pleasantly surprised. The release date is similarly fuzzy: the only snapshot is dated July 15 and OpenRouter recorded the listing at 22:16 UTC on July 27. I could not find the changelog entry several aggregators cite as July 25.
Go and count your prompts
Measure your input length distribution before you do anything else. Not the mean, the histogram, and specifically the share of requests sitting above 32,000 tokens. That single number decides whether this model is a fourfold saving or a mild regression, and no amount of blended-rate arithmetic will tell you which.
If the mass of your traffic sits well under 32K, this is the cheapest way to buy tokens right now and the missing benchmarks matter less, because on short prompts you can evaluate quality yourself in an afternoon on your own data. If you are straddling the boundary, the fix is boring and effective: trim the prompt to stay under it, or accept the second tier and then ask honestly why you would run an unmeasured model at DeepSeek's exact price. If your prompts are routinely six figures, this is not your model. A flat card beats a stepped one every time your inputs are long, and DeepSeek has both the flat card and a published score.
I would still run my own evaluation before ruling it out. A $0.03 rate is a genuinely different price point and it deserves a look, and on short prompts the trial costs almost nothing. Just build the test set around the prompts you actually send, and check the token counts before you check the outputs.
Sources
- Alibaba Cloud Model Studio: Qwen3.7 Flash - The four regional CNY price tables headed by input length, the 1M context with 991,808 usable input and 65,536 max output, the 983,616 and 262,144 thinking-mode limits, the cache percentages, the batch-file discount marked Beijing-only, and the 5M TPM and 30,000 versus 15,000 RPM rate limits
- QwenCloud: Qwen3.7 Flash model page - The rendered USD tier tables, the July 15 snapshot note, and the 131K max-output figure that conflicts with both the Model Studio doc and OpenRouter
- OpenRouter: models API - Independent confirmation of the tier boundaries at exactly 32,000 and 256,000 prompt tokens, the second and third tier rates, the 0.20x cache-read and 1.25x cache-write multipliers, 65,536 max completion tokens, a single Alibaba first-party endpoint, and the 2026-07-27 listing timestamp
- DeepSeek: API pricing - $0.14 input, $0.0028 on a cache hit and $0.28 output, flat across the full 1M context, which is the card that ties Qwen's middle tier at $0.175 blended and beats it above 256K
- Artificial Analysis: model comparison data - The Intelligence Index figures for DeepSeek V4 Flash at 49.93, GPT-5.6 Luna at 51.24 on max effort, Qwen3.7 Plus at 38.98, Gemini 3.5 Flash-Lite at 36.48 and Mistral Small 4 at 19.56, and the absence of any Qwen3.7 Flash entry from the model list, the leaderboard payload and the Alibaba provider page
- OpenAI: API pricing - GPT-5.6 Luna at $0.20 and $1.20 after the July 30 cut, plus the 50% batch discount that Qwen's international endpoints do not offer
- Google: Gemini API pricing - Gemini 3.5 Flash-Lite at $0.30 and $2.50 with 50% off for batch and no context-length tiering anywhere in the 1M window
- Mistral: pricing - Mistral Small 4 at $0.15 and $0.60 on a flat Apache-licensed card, and the 256K context limit that keeps it out of the repo-pass row entirely
- Hugging Face: models API - Zero results for Qwen3.7 under the Qwen organisation, confirming the generation is closed and that no third party can run an independent evaluation