Ling-3.0-flash sells for $0.021, $0.06 and $0.075 per million input tokens right now. Same weights, same week, and the identical benchmark run bills $20.70 or $72.71 depending only on who you buy it from.
Ant Group opened the weights to a 124B mixture-of-experts model on August 2 under MIT, ten days after quietly putting it on the API and one day before the free trial ran out. The model is genuinely interesting: 5.1B parameters active per token, 256K of native context, and the highest intelligence score Artificial Analysis has measured in the price band it competes in. The pricing is where it gets strange. Novita and Vercel list it at $0.06 and $0.18. OpenRouter and ZenMux both sell at $0.021 and $0.063, a flat 65% off with no expiry printed anywhere, and they reach the model through different suppliers. DeepInfra charges $0.075 and $0.22 and hands you half the context for the privilege. Artificial Analysis priced its benchmark run at that dearest card, which is why its headline cost sits at $72.70. It also published the token counts, so we repriced the run ourselves: $59.13 in the middle, $20.70 on the promo, same model and same score. And then there is the part that undoes most of the advantage. Ling knocks 80% off a cached token. DeepSeek knocks off 98%. Run a cache-heavy job and DeepSeek's far worse sticker lands within a tenth of a cent of Ling's, then goes ahead of it entirely once your hit rate clears 84.2%.

Photo by Kier in Sight Archives on Unsplash
Read this before you pick a host
A good cheap model with a pricing situation nobody has written down properly. Three hosts, three cards, a 3.5x spread on identical work, and a promotional rate that could vanish without notice because no document commits to keeping it. Buy it at $0.021 and it is the cheapest credible model we can currently find. Budget it at $0.021 and you have a problem waiting.
Spread on one rate card
3.5x
$0.021 to $0.075 input
Intelligence Index
37.82
1st of 56 in its price class
Output tokens per task
35,774
74% of it reasoning
Three rate cards, one set of weights
Start with the thing that makes this model hard to write about. There is no single price for Ling-3.0-flash. We found three distinct price points across at least four hosts, and they are not a case of resellers disagreeing about one wholesale rate.
| Where you buy it | Input / 1M | Output / 1M | Cache read | Context |
|---|---|---|---|---|
| OpenRouter / ZenMux | $0.021 | $0.063 | $0.0042 | 262,144 |
| Novita, Vercel AI Gateway | $0.06 | $0.18 | $0.012 | 262,144 |
| DeepInfra, and InclusionAI per AA | $0.075 | $0.22 | $0.015 | 131,072 |
The promotional card is arithmetically exact against the list one. $0.06 times 0.35 is $0.021, $0.18 times 0.35 is $0.063, $0.012 times 0.35 is $0.0042. OpenRouter labels it a 65% discount and ZenMux advertises the same 65%. What matters is where the money comes from: Novita's own API reports its origin price and its current price as the same figure, so Novita is not discounting anything. The routers are absorbing it. That is a marketing spend, and marketing spends end.
Worth knowing that the two cheap routers do not even reach the model the same way. OpenRouter fronts Novita. ZenMux buys from Ant Ling directly. They arrive at an identical card from different suppliers, which is a hint that the 65% is a number both chose rather than one they were given.
We looked for an end date and could not find one on either router, in Ant's press release, or on the model card. Treat the $0.021 as a price you are enjoying rather than a price you are owed.
DeepInfra is the odd one out and the worse deal twice over. It costs 3.57x the promotional input rate, and its deployment caps context at 131,072 tokens against the 262,144 the model was actually trained to handle. If you picked a host by name recognition and landed there, you are paying more for less of the model. The one thing in its defence: Artificial Analysis records that same $0.075 card as InclusionAI's own first-party rate, so DeepInfra may simply be the host not discounting below what Ant charges.
What Ant shipped, and the order it shipped in
The sequencing here is deliberate and worth noticing. The API went live on July 23. Ant's English press release followed on July 27 and promised free access through August 3, after which, in its own words, the weights would be open-sourced. The BF16 checkpoint landed on Hugging Face on August 2 at 16:14 UTC. The quantised checkpoints followed on August 4.
So the model ran as a free hosted API for ten days, collected whatever it collected, and only then became something you could run yourself. A good chunk of the SEO coverage still sitting on page one of search was written during that window and says this model is API-only with no weights and no licence. That was true. It stopped being true five days ago.
| Spec | Value |
|---|---|
| Parameters | 124B total, 5.1B active per token |
| Experts | 512 routed plus 1 shared, top-8 activated |
| Attention | 35 KDA layers and 7 gated MLA layers, 5:1 |
| Context | 262,144 native, 32,768 max output |
| Licence | MIT, ungated, all four checkpoints |
| Reasoning | On by default, opt out per request |
One correction worth making loudly, because it is circulating. Some write-ups say 51B active parameters. It is 5.1B. Ant states it twice in primary sources, and the ratio it quotes against its own 1T-class flagship, 8.1% of Ring-2.6-1T's 63B active, only works with the decimal in the right place. A tenfold error in active parameters is the difference between a model that fits on one GPU and one that does not.
The 1M context claim in Ant's materials is a different matter. The config file says 262,144, and every host that actually serves the model caps at 262,144 or below. Read the million as an aspiration about rope scaling, not a number you can send tokens to today.
The same benchmark run, billed three ways
Artificial Analysis publishes a cost to run its Intelligence Index for each model, and for Ling-3.0-flash that figure is $72.70. It also publishes the cost broken into five components, which is generous of it, because those components divide back into token counts. AA's price fields for this model are $0.075, $0.22 and $0.015. It records those as InclusionAI's first-party rate, and they happen to be exactly what DeepInfra charges too. Either way it is the dearest of the three cards in the market.
Dividing each cost line by the rate it was charged at gives roughly 240M output tokens, 542M cache reads, 142M cache writes and 14M fresh input tokens. Recharging those identical volumes at the other two cards is then simple arithmetic.
One check before trusting any of that. Multiplying our back-solved tokens by the rates we divided by would prove nothing, since it is the same operation run backwards. The real test is that AA separately publishes a canonical input token count for the run, and our three input legs sum to 697,155,733 against its published 697,155,771. That is a gap of 38 tokens in seven hundred million, so the volumes are right.
| Host | Cost to run the index | Per task | Score |
|---|---|---|---|
| At $0.021 / $0.063 (OpenRouter, ZenMux) | $20.70 | $0.0107 | 37.82 |
| At $0.06 / $0.18 (Novita, Vercel) | $59.13 | $0.0306 | 37.82 |
| At $0.075 / $0.22 (what AA used) | $72.71 | $0.0376 | 37.82 |
Three rows, one model, one score, and the bill moves 3.5x. Every public leaderboard that ranks models by cost is implicitly ranking them by whichever host the benchmarker happened to call, and almost none of them say which one that was. Two data-quality notes on the AA page while you are there: it lists Ling-3.0-flash as proprietary, which stopped being true on August 2, and dates it to August 2026 when the API opened in July. The 56-model class it tops is a price band, not an open-weights cohort, so first of 56 means first among cheap models generally rather than first among models you can download.
It thinks for 26,442 tokens before it answers
Per benchmark task, Ling-3.0-flash emits 35,774 output tokens. Of those, 9,331 are the answer and 26,442 are reasoning. Three quarters of what you pay for on the output leg is the model talking to itself.
Across the whole index that is 240M output tokens against a class median of 63M, so this model is 3.8x more verbose than its peers and AA flags it as very verbose. Reasoning is on by default in the chat template. You can pass a flag to switch it off, and if your workload is extraction or classification rather than problem solving, you probably should, because the 74% of output you are not reading is still metered.
This is the mechanism that keeps eating cheap models. A rate card 6.7x below DeepSeek's does not produce a bill 6.7x smaller if the model writes several times as much to reach the same place. It is the same trap we walked through on Qwen3.8-Max, where a 60% output discount turned into an 11% smaller invoice.
The cache discount decides this, not the sticker
Here is the finding that changed how we read this launch. Ling takes 80% off a cached input token on all three hosts. That sounds fine until you notice the industry has moved past it. DeepSeek takes 98% off. Anthropic and OpenAI take 90%.
ZenMux publishes a daily cache hit rate for Ling-3.0-flash, and on August 6 it read 82.84%. That is real traffic rather than an assumption we picked, but it is a daily figure and it moves, so treat it as a snapshot. Blend each model's input leg at that rate and the ordering changes.
| Model | Miss | Hit | Discount | Blended at 82.84% |
|---|---|---|---|---|
| Ling-3.0-flash (promo) | $0.021 | $0.0042 | 80% | $0.00708 |
| Qwen3.7 Flash, up to 32K | $0.03 | $0.006 | 80% | $0.01012 |
| Ling-3.0-flash (list) | $0.06 | $0.012 | 80% | $0.02024 |
| GLM-4.7-FlashX | $0.07 | $0.01 | 86% | $0.02030 |
| Ling-3.0-flash (DeepInfra) | $0.075 | $0.015 | 80% | $0.02530 |
| DeepSeek V4 Flash 0731 | $0.14 | $0.0028 | 98% | $0.02634 |
| GPT-5.6 Luna | $0.20 | $0.02 | 90% | $0.05089 |
Look at the two rows in the middle of that list. DeepSeek's sticker is 1.87x Ling's on DeepInfra. After caching, the two land at $0.02634 and $0.02530, so DeepSeek is only 4.1% dearer. Its 98% cache discount has eaten nearly all of a sticker price that looks almost twice as expensive, and it does it on precisely the workload people buy cheap models for, which is agents replaying a long stable prefix.
Push the hit rate a little higher and it flips outright. The two cards cross at 84.2%: below that Ling is cheaper on input, above it DeepSeek is. At 90% cache hits DeepSeek costs $0.01652 against Ling's $0.02100, so the model with double the sticker price is 21% cheaper to feed. Which of these two is the budget option genuinely depends on how well your prefixes cache, and that is not a question a rate card can answer for you.
Two caveats we owe you on that comparison. We have given Ling its cheapest reseller price and DeepSeek only its first-party one; on OpenRouter, DeepSeek V4 Flash sells for $0.09 and $0.18 with a shallower 80% cache discount. And DeepSeek's own pricing page currently carries a warning that a significant increase is planned, with no date attached. Neither changes the shape of the argument, but both cut against it.
Then check what you get for the money. DeepSeek V4 Flash scores 51.77 on the same index where Ling scores 37.82, and AA puts its cost per task at $0.0271 against Ling's $0.0376. On the pricing AA used, DeepSeek is both cheaper per finished task and 14 index points better. Ling only wins the cost argument at $0.021, where it drops to $0.0107 per task, roughly 40% of what DeepSeek charges to finish the same work. The entire case for this model rests on a discount nobody has promised to keep.
77 GB is the most interesting number here
Activating 5.1B of 124B parameters, roughly 1 in 24 by weight, does something useful to the download. Ant counts its sparsity differently and quotes 1 in 64, which is the expert ratio, 8 routed experts of 512. Both are true and they are not the same number. Either way, Ant shipped four checkpoints and two of them fit on a single 80GB card.
| Checkpoint | Size | Fits on |
|---|---|---|
| BF16 | 255.00 GB | 4x H100 80GB |
| FP8 | 128.47 GB | 2x H100, or 1x H200 141GB |
| int4 | 77.04 GB | 1x H100 80GB, just |
| FP4 | 70.40 GB | 1x H100 80GB |
A 124B model on one H100 is a genuinely different proposition from the 1T-class open weights everyone spent the spring failing to run. Mind the units when you plan it, though. An 80GB H100 is 80 GiB, which is about 85.9 GB, and the 77.04 GB int4 checkpoint is 71.75 GiB. That leaves roughly 8.25 GiB for the KV cache and activations, which is workable at modest context but nowhere near enough to serve 256K. FP4 leaves about 14.4 GiB and is the more comfortable single-card option.
Two practical warnings. The FP4 repository declares its quantization method as fp8 in the config and carries an 8-bit tag, so check what you actually pulled before you size a cluster around it. And vLLM support is not upstream: you need Ant's own fork. SGLang's cookbook page is the better-supported path, with a dedicated image and a tested tensor-parallel recipe. If you are weighing this against renting, our Kimi K3 self-hosting breakdown has the GPU-hour arithmetic that applies here too.
Five of Ant's headline scores are pixels in a bar chart
Ant does publish machine-readable evaluation results, and credit where it is due, they are structured properly in the repository rather than buried in a README table. Every one is self-reported.
| Benchmark | Ant's score | Who ran it |
|---|---|---|
| AIME 2026 | 93.2 | Ant |
| HMMT Feb 2026 | 87.0 | Ant |
| SWE-bench Multilingual | 72.4 | Ant, under OpenHands |
| SWE-bench Pro | 56.6 | Ant, under OpenHands |
| Humanity's Last Exam | 22.7 | Ant |
| Intelligence Index | 37.82 | Artificial Analysis |
What is missing is more interesting than what is there. The README names Terminal-Bench 2.1, GDPval v2, MCP-Atlas, SkillsBench, MiniAppBench and several others, and describes in detail how each was run. The scores themselves appear nowhere as text. They exist only inside two images served from an Alipay domain, one a grid of bar charts and one a rendered table. The numbers are printed legibly on both, so this is not a case of squinting at bar heights. It is that none of it is machine-readable. You cannot copy a value, sort the set, diff it against a later release, or drop it into a spreadsheet beside anybody else's run without retyping it by hand and hoping you did not fat-finger a digit.
There is also no technical report and no paper. Previous generations got a Ling-V2 repository; this one has only a vLLM fork. For a model released under MIT with this much engineering in it, that is a strange gap.
Plan for the price that will still be there
Our read, stated plainly. Ling-3.0-flash at $0.021 and $0.063 is the best price per unit of capability we can currently find, and if you have batch work that tolerates a 37.82 rather than a 51.77, routing it through OpenRouter or ZenMux today is an easy win. Take the win. Just do not build a forecast on it.
The rate that will still exist in six months is $0.06 and $0.18, because that is the one Novita charges for the compute. Model your spend there. If the promo holds, you are pleasantly wrong by 65%. If it ends, nothing breaks. Anyone who wrote $0.021 into a budget spreadsheet watches their input line nearly triple the morning a router decides it has bought enough traffic.
Two things would change our mind about the model itself. An independent measurement of the benchmarks currently locked in those PNGs, and a cache discount that matches what DeepSeek and OpenAI already offer. The second is the one that would actually move bills, because on agentic traffic the cached leg is most of the input, and at 80% off Ling is tied with Qwen3.7 Flash for the shallowest cache discount in the table above. Since the crossover against DeepSeek sits at an 84.2% hit rate, your own cache behaviour decides this more than either rate card does. You can check that against your traffic mix in the cost calculator or line the cards up side by side on the pricing page.
Sources
- Hugging Face: inclusionAI/Ling-3.0-flash - The MIT licence on all four checkpoints, the 124B total and 5.1B active figures, the 35 KDA plus 7 gated MLA layer split, 512 routed experts plus 1 shared with top-8 routing, the 262,144 value of max_position_embeddings, thinking enabled by default in the chat template, and the August 2 BF16 upload followed by the August 4 quantised checkpoints. Checkpoint sizes were summed from the blob metadata: 255.00 GB, 128.47 GB, 77.04 GB and 70.40 GB
- Artificial Analysis: Ling-3.0-flash - The Intelligence Index of 37.82, first of 56 in its class against a class median of 8, the $72.70 cost to run the index and its five-way split of $48.53 reasoning output, $4.39 answer output, $10.64 cache write, $8.13 cache read and $1.02 fresh input, the 240M total output tokens against a 63M class median, the 35,774 output tokens per task split 9,331 answer and 26,442 reasoning, the $0.03757 cost per task, and the $0.075 / $0.22 / $0.015 price fields, which AA attributes to InclusionAI as a first-party source and which also match DeepInfra's public card exactly. The canonical input token count of 697,155,771 is what our back-solved volumes were checked against. Also the comparison figures for DeepSeek V4 Flash 0731 at 51.77 and $0.0271 per task
- OpenRouter: inclusionai/ling-3.0-flash - The $0.021 / $0.063 / $0.0042 card carrying an explicit 0.65 discount field on the Novita endpoint, the separate DeepInfra endpoint at $0.075 / $0.22 / $0.015 with a 131,072 context against Novita's 262,144, the July 23 creation timestamp, and the free variant that now returns zero endpoints
- Novita: Ling-3.0-flash - The $0.06 / $0.18 / $0.012 list card, and the API field reporting origin price and current price as identical, which is what establishes that the 65% discount is funded by the routers rather than by the host
- ZenMux: Ling-3.0-flash - The matching 65% off card, the July 23 publish date, the fact that ZenMux sources from Ant Ling directly rather than through Novita, and the 82.84% daily cache hit rate for August 6, 2026 used for every blended input figure in this post. That figure is a daily reading and will have moved by the time you read this
- OpenRouter: Qwen3.7 Flash - $0.03 input, $0.006 cached and $0.13 output. The $0.03 applies only up to 32K tokens: Alibaba tiers it to $0.10 between 32K and 256K and $0.20 above that, so the blended figure in the table above is a best case rather than a flat rate
- Ant Group press release, July 27, 2026 - The 124B and 5.1B figures repeated in Ant's own words, the free API through August 3 followed by open weights, the claim of matching models at two to three times the parameter scale, and the planning-execution framing that positions this as an execution node rather than a frontier model. Note this is a paid press release rather than reported coverage
- SGLang: Ling-3.0-flash cookbook - The official serving recipe, the dedicated Docker image, and the tensor-parallel guidance behind the per-card fit figures. vLLM support requires Ant's vllm-ling-v3 fork rather than upstream
- DeepSeek: API pricing - V4 Flash at $0.14 on a cache miss and $0.0028 on a hit, the 98% discount that does most of the work in the blended table above, and $0.28 output. The same page carries DeepSeek's notice that it plans a significant price increase, with no effective date given
- OpenAI: API pricing - GPT-5.6 Luna at $0.20 input, $0.02 cached and $1.20 output, the 90% cache discount that is closer to the current norm than Ling's 80%
- Z.ai: pricing - GLM-4.7-FlashX at $0.07 input, $0.01 cached and $0.40 output, the third 80-something percent discount in the comparison table