TokenBlog
Model releases, pricing breakdowns, and practical guides for developers.
Grok 4.6 costs exactly what Grok 4.5 costs on input, on output, and at the cliff. The only…
xAI shipped Grok 4.6 on August 12 and put it on the rate card of the model it succeeds. Input $2.00, output $6.00, a 500,000 token window, and the same doubling at 200,000 prompt tokens. Four lines…
Google cut Gemini 3.6 Flash in half on Thursday and launched Gemini 3.7 Flash at the same…
Gemini 3.7 Flash went GA on August 13 at $0.75 input and $3.75 output per million tokens, and the coverage all repeated the same framing…
12 min readMeta gave Muse Glimmer's weights away yesterday and exactly one company sells it. $0.35…
Meta shipped a 30B dense model under Apache 2.0 on August 10, its first open weights in more than a year, and named twelve launch partners…
11 min readImagen 4 stops answering on August 17. Google's pages send you to two different…
Three endpoints go dark next Monday, and the retirement changes the billing unit rather than just the number. Imagen 4 sold an image for a…
13 min readOn August 1 we wrote that nobody was going to bill you $20 per million input tokens. Four…
OpenAI's changelog for August 5 is one sentence long and mentions no price: Fast mode now supports long-context requests for GPT-5.6 Sol…
11 min readAnthropic will now stop an agent session at a number you choose, written in whole cents as…
Session budgets shipped for Claude Managed Agents on August 7. Set budget.max_list_cost.amount to "2500" and the session stops at $25, which…
12 min readDeepSeek says a significant price increase is coming and will not say how significant. We…
A three-sentence notice went up on DeepSeek's pricing page on August 6, as footnote number two under the rate table. No percentage, no…
11 min readLing-3.0-flash sells for $0.021, $0.06 and $0.075 per million input tokens right now. Same…
Ant Group put a 124B MoE on the API on July 23 and did not open the weights until August 2, one day before the free trial expired. What…
12 min readMeta will cut your API bill 22x for permission to train on your prompts. The rate that…
Muse Spark 1.2 shipped on August 5 with a second model ID beside it. muse-spark-1.2-contributor bills $0.10 input, $0.002 cached input and…
11 min readgrok-voice-latest moved to Think Fast 2.0 this morning, so unchanged code now bills 60%…
xAI repointed the grok-voice-latest alias today, taking anyone who used it from $0.05 a minute of audio to $0.08. That is $3.00 an hour…
13 min readQwen3.8-Max prices output 60% under Kimi K3. On the same benchmark suite the bill came in…
Alibaba took its 2.4T-parameter flagship to general availability on August 3 and every gateway now quotes $2 input and $6 output per…
11 min readQwen3.7 Flash lists at $0.03 a million. One token past 32K, the identical request bills…
Alibaba priced its newest cheap tier at $0.03 input and $0.13 output, which is the lowest sticker on any 1M-context model you can call…
10 min readDeepSeek kept the price at $0.14 and added ten index points. Eleven models now tie on…
On July 31 DeepSeek pushed a new build behind an endpoint that already existed, kept the model ID, kept the architecture and kept every…
10 min readOpenAI and Anthropic both charge exactly double for speed. One of them tells you how much…
Latency stopped being a capacity commitment and became a per-request purchase in the space of eight days. Anthropic closed Priority Tier to…
11 min readOpenAI cut GPT-5.6 Luna by 80% and left Sol untouched. That is a 25x spread inside one…
On July 30 Luna went from $1.00 and $6.00 per million tokens to $0.20 and $1.20, an identical 80% off input, output and cached input. Terra…
9 min readAnthropic says one cryptographic attack cost it about $100,000 in API time. Its own price…
Labs almost never publish what a task cost them. On July 28 Anthropic did, twice: roughly $100,000 per result, one billion output tokens…
9 min readClaude Sonnet 5's $2 and $10 promo ends on August 31. Every line on the rate card moves up…
Anthropic disclosed the expiry in the launch post on June 30, then encoded it as a second row in the docs, so September 1 is a promotion…
10 min readKimi K3's weights went up for free on Monday and every host still matches Moonshot's…
Open weights normally start a price war. Moonshot published the full Kimi K3 checkpoint on July 27 and ten providers now serve it, none of…
11 min readPoolside is selling a 118B coding model for 10 cents in and 20 cents out. It also…
Laguna S 2.1 shipped July 21 with open weights and a rate card of $0.10 input and $0.20 output per million tokens, a fiftieth of GPT-5.6 Sol…
10 min readGoogle's Lite tier keeps getting less lite. A year of releases took Flash-Lite output from…
Gemini 3.5 Flash-Lite shipped July 21 at $0.30 input and $2.50 output per million tokens, against $0.25 and $1.50 for the model it replaces…
9 min readDeepSeek switched off deepseek-chat and deepseek-reasoner yesterday. The rename is one…
On July 24 at 15:59 UTC, DeepSeek retired its two oldest API model names, and requests to deepseek-chat and deepseek-reasoner now fail…
8 min readAnthropic shipped Opus 5 at the same $5/$25 it charged for 4.8. The story is what that…
Claude Opus 5 landed July 24 with a rate card copied from Opus 4.8, $5 input and $25 output per million tokens, unchanged for a fourth Opus…
8 min readEvery big lab now ships a dedicated coding model. Their output prices run from $0.80 to…
Once you filter the July 2026 rate cards down to the models sold specifically for coding agents, the label stops meaning anything about…
8 min readClaude Opus 4.8 just slipped off the LLM value frontier. A $3 model scores higher on the…
Rank every current model by Artificial Analysis's Intelligence Index and by what it actually costs to run that eval suite, and the July 2026…
9 min readAlibaba says Qwen3.8 Max is the world's second-best model. It won't publish a benchmark…
Qwen3.8 Max previewed July 19 at WAIC Shanghai, two days after Kimi K3 went open weights, and Alibaba called it second only to Claude Fable…
7 min read






















































































































































