Skip to main content
TokenCost logoTokenCost

Ling-3.0-flash

Last verified August 7, 2026 · Ant Group pricing

$0.060/1M input · $0.180/1M output · 262K context · Ant Group

Count Tokens

Tokens
Input Cost
Output Cost

Estimate Monthly Cost

Monthly Cost Estimator

Quick:
<$0.0001/mo

Pick a preset above or enter custom usage

Alternatives to Ling-3.0-flash

Pricing Details

Ant GroupTHREE DISTINCT PRICES across at least four hosts. The $0.06/$0.18 listed here is the middle card, quoted by both Novita (which serves the model, and whose API reports origin price and current price as identical) and the Vercel AI Gateway. OpenRouter and ZenMux both sell at $0.021/$0.063 (cache read $0.0042), an exact 65% discount funded by the routers with no published end date - do not budget against it. They reach the model differently: OpenRouter fronts Novita, ZenMux buys from Ant Ling direct. DeepInfra runs its own deployment at $0.075/$0.22 (cache read $0.015) and caps context at 131,072 rather than 262,144; Artificial Analysis records that same card as InclusionAI's first-party rate. Ant's own console is Alipay-gated and publishes no English rate card. Cache read is 20% of input on every host, an 80% discount against DeepSeek's 98% and the 90% OpenAI and Anthropic offer, so on cache-heavy agentic traffic at the 82.84% hit rate ZenMux reported on 2026-08-06, Ling on DeepInfra blends to $0.02530/1M input against DeepSeek V4-Flash's $0.02634 despite a sticker that looks 1.87x cheaper - and the two cross over at an 84.2% hit rate, above which DeepSeek is outright cheaper on input. 124B total params with 5.1B active (some coverage says 51B, which is wrong by 10x); 512 routed experts plus 1 shared, top-8 activated; 42 layers of hybrid attention, 35 KDA to 7 gated MLA in a 5:1 pattern. Reasoning is ON by default and must be disabled per request. API launched 2026-07-23, free tier ran through 2026-08-03, BF16 weights opened 2026-08-02 under MIT, quantised checkpoints 2026-08-04 (BF16 255.00 GB, FP8 128.47 GB, int4 77.04 GB, FP4 70.40 GB, so int4 and FP4 fit one 80GB H100). Ant claims 1M context but max_position_embeddings is 262,144 and no host serves more. Artificial Analysis Intelligence Index 37.82, 1st of 56 in its price class against a class median of 8, but very verbose: 240M output tokens to run the index against a 63M class median, and 35,774 output tokens per task of which 26,442 are reasoning. AA priced its run at the $0.075/$0.22/$0.015 card, so its published $72.70 cost to run the index is the most expensive of the three; the same measured token volumes cost $59.13 at $0.06/$0.18 and $20.70 at $0.021/$0.063. Self-reported by Ant: AIME 2026 93.2, HMMT Feb 2026 87.0, SWE-bench Multilingual 72.4, SWE-bench Pro 56.6, Humanity's Last Exam 22.7. Several other benchmarks named in Ant's README (Terminal-Bench 2.1, GDPval v2, MCP-Atlas, SkillsBench, MiniAppBench) are published only inside images, legible but not machine-readable. No technical report. SGLang has an official cookbook and image; vLLM needs Ant's vllm-ling-v3 fork rather than upstream.
Input / 1M tokens
$0.06
Output / 1M tokens
$0.18
Context Window
262K
Max Output
33K

Price History

Launched at current price on 2026-07-23. No price changes recorded.

Frequently Asked Questions

Common questions about Ling-3.0-flash pricing and usage

Read More About Ling-3.0-flash