Skip to main content
TokenCost logoTokenCost

TokenBlog

Model releases, pricing breakdowns, and practical guides for developers.

Model ReleaseAugust 15, 2026
Grok
Latest·11 min read

Grok 4.6 costs exactly what Grok 4.5 costs on input, on output, and at the cliff. The only…

xAI shipped Grok 4.6 on August 12 and put it on the rate card of the model it succeeds. Input $2.00, output $6.00, a 500,000 token window, and the same doubling at 200,000 prompt tokens. Four lines…

Model ReleaseAugust 14, 2026
Gemini

Google cut Gemini 3.6 Flash in half on Thursday and launched Gemini 3.7 Flash at the same…

Gemini 3.7 Flash went GA on August 13 at $0.75 input and $3.75 output per million tokens, and the coverage all repeated the same framing…

12 min read
Model ReleaseAugust 11, 2026
Model Release

Meta gave Muse Glimmer's weights away yesterday and exactly one company sells it. $0.35…

Meta shipped a 30B dense model under Apache 2.0 on August 10, its first open weights in more than a year, and named twelve launch partners…

11 min read
GuideAugust 10, 2026
Guide

Imagen 4 stops answering on August 17. Google's pages send you to two different…

Three endpoints go dark next Monday, and the retirement changes the billing unit rather than just the number. Imagen 4 sold an image for a…

13 min read
IndustryAugust 9, 2026
GPT

On August 1 we wrote that nobody was going to bill you $20 per million input tokens. Four…

OpenAI's changelog for August 5 is one sentence long and mentions no price: Fast mode now supports long-context requests for GPT-5.6 Sol…

11 min read
GuideAugust 8, 2026
Claude

Anthropic will now stop an agent session at a number you choose, written in whole cents as…

Session budgets shipped for Claude Managed Agents on August 7. Set budget.max_list_cost.amount to "2500" and the session stops at $25, which…

12 min read
IndustryAugust 8, 2026
DeepSeek

DeepSeek says a significant price increase is coming and will not say how significant. We…

A three-sentence notice went up on DeepSeek's pricing page on August 6, as footnote number two under the rate table. No percentage, no…

11 min read
Model ReleaseAugust 7, 2026
Model Release

Ling-3.0-flash sells for $0.021, $0.06 and $0.075 per million input tokens right now. Same…

Ant Group put a 124B MoE on the API on July 23 and did not open the weights until August 2, one day before the free trial expired. What…

12 min read
IndustryAugust 6, 2026
Industry

Meta will cut your API bill 22x for permission to train on your prompts. The rate that…

Muse Spark 1.2 shipped on August 5 with a second model ID beside it. muse-spark-1.2-contributor bills $0.10 input, $0.002 cached input and…

11 min read
Model ReleaseAugust 5, 2026
Grok

grok-voice-latest moved to Think Fast 2.0 this morning, so unchanged code now bills 60%…

xAI repointed the grok-voice-latest alias today, taking anyone who used it from $0.05 a minute of audio to $0.08. That is $3.00 an hour…

13 min read
Model ReleaseAugust 4, 2026
Qwen

Qwen3.8-Max prices output 60% under Kimi K3. On the same benchmark suite the bill came in…

Alibaba took its 2.4T-parameter flagship to general availability on August 3 and every gateway now quotes $2 input and $6 output per…

11 min read
Model ReleaseAugust 3, 2026
Qwen

Qwen3.7 Flash lists at $0.03 a million. One token past 32K, the identical request bills…

Alibaba priced its newest cheap tier at $0.03 input and $0.13 output, which is the lowest sticker on any 1M-context model you can call…

10 min read
Model ReleaseAugust 2, 2026
DeepSeek

DeepSeek kept the price at $0.14 and added ten index points. Eleven models now tie on…

On July 31 DeepSeek pushed a new build behind an endpoint that already existed, kept the model ID, kept the architecture and kept every…

10 min read
IndustryAugust 1, 2026
Industry

OpenAI and Anthropic both charge exactly double for speed. One of them tells you how much…

Latency stopped being a capacity commitment and became a per-request purchase in the space of eight days. Anthropic closed Priority Tier to…

11 min read
IndustryJuly 31, 2026
GPT

OpenAI cut GPT-5.6 Luna by 80% and left Sol untouched. That is a 25x spread inside one…

On July 30 Luna went from $1.00 and $6.00 per million tokens to $0.20 and $1.20, an identical 80% off input, output and cached input. Terra…

9 min read
ResearchJuly 30, 2026
Anthropic

Anthropic says one cryptographic attack cost it about $100,000 in API time. Its own price…

Labs almost never publish what a task cost them. On July 28 Anthropic did, twice: roughly $100,000 per result, one billion output tokens…

9 min read
IndustryJuly 29, 2026
Claude

Claude Sonnet 5's $2 and $10 promo ends on August 31. Every line on the rate card moves up…

Anthropic disclosed the expiry in the launch post on June 30, then encoded it as a second row in the docs, so September 1 is a promotion…

10 min read
ResearchJuly 28, 2026
Kimi

Kimi K3's weights went up for free on Monday and every host still matches Moonshot's…

Open weights normally start a price war. Moonshot published the full Kimi K3 checkpoint on July 27 and ten providers now serve it, none of…

11 min read
Model ReleaseJuly 27, 2026
Model Release

Poolside is selling a 118B coding model for 10 cents in and 20 cents out. It also…

Laguna S 2.1 shipped July 21 with open weights and a rate card of $0.10 input and $0.20 output per million tokens, a fiftieth of GPT-5.6 Sol…

10 min read
ResearchJuly 26, 2026
Gemini

Google's Lite tier keeps getting less lite. A year of releases took Flash-Lite output from…

Gemini 3.5 Flash-Lite shipped July 21 at $0.30 input and $2.50 output per million tokens, against $0.25 and $1.50 for the model it replaces…

9 min read
GuideJuly 25, 2026
DeepSeek

DeepSeek switched off deepseek-chat and deepseek-reasoner yesterday. The rename is one…

On July 24 at 15:59 UTC, DeepSeek retired its two oldest API model names, and requests to deepseek-chat and deepseek-reasoner now fail…

8 min read
Model ReleaseJuly 25, 2026
Claude

Anthropic shipped Opus 5 at the same $5/$25 it charged for 4.8. The story is what that…

Claude Opus 5 landed July 24 with a rate card copied from Opus 4.8, $5 input and $25 output per million tokens, unchanged for a fourth Opus…

8 min read
ComparisonJuly 24, 2026
Comparison

Every big lab now ships a dedicated coding model. Their output prices run from $0.80 to…

Once you filter the July 2026 rate cards down to the models sold specifically for coding agents, the label stops meaning anything about…

8 min read
ResearchJuly 23, 2026
Research

Claude Opus 4.8 just slipped off the LLM value frontier. A $3 model scores higher on the…

Rank every current model by Artificial Analysis's Intelligence Index and by what it actually costs to run that eval suite, and the July 2026…

9 min read
Model ReleaseJuly 22, 2026
Qwen

Alibaba says Qwen3.8 Max is the world's second-best model. It won't publish a benchmark…

Qwen3.8 Max previewed July 19 at WAIC Shanghai, two days after Kimi K3 went open weights, and Alibaba called it second only to Claude Fable…

7 min read