Skip to main content
TokenCost logoTokenCost
Model ReleaseJuly 20, 2026·7 min read

Mira Murati's lab open-sourced a frontier-class model. The strange part is how hard it is to just buy tokens for it.

Inkling arrived on July 15 with a 975-billion-parameter Mixture-of-Experts, an Apache 2.0 license, and an Artificial Analysis score that tops every other US open-weights model. What it does not have is a first-party API. Five days in, there is exactly one place to pay per token, and the number everyone keeps quoting is not even an inference rate.

Dark grid of geometric cells representing tokens and API cost for Inkling

Photo by Samuel Scalzo on Unsplash

The situation in a paragraph: the weights are free and the capability is real, but the pricing is a mess of its own making. Thinking Machines shipped Inkling as open weights and pointed everyone at partners to actually serve it, but almost none of those partners publish a per-token rate. The clean number is OpenRouter's $1 in and $4.05 out per million tokens, from one lone provider. The $1.87/$4.68 you will see repeated on aggregators is the Tinker fine-tuning meter wearing an API costume. So the honest headline is not "Inkling is cheap" or "Inkling is expensive." It is "Inkling barely has a price yet," and if you are budgeting around it, that is the thing to plan for.

The model, in one breath

Inkling is the first open model out of Thinking Machines Lab, the company Mira Murati started after leaving OpenAI. It is a Mixture-of-Experts with 975 billion total parameters and 41 billion active per token, 256 routed experts plus two shared, a 1 million token context window, and a 45-trillion-token training run across text, image, audio, and video. Input is multimodal; output is text only, so despite a strong VoiceBench score it does not speak back. The license is Apache 2.0, and the weights, including a 4-bit NVFP4 checkpoint, sit on Hugging Face.

The scores are the reason anyone is paying attention. Artificial Analysis, the one group that has measured Inkling independently rather than reprinting the model card, puts it at 41 on its Intelligence Index and calls it the leading US open-weights model. Thinking Machines' own card reports 77.6% on SWE-bench Verified, 87.2% on GPQA Diamond, and 97.1% on AIME 2026. Take the self-reported figures with the usual pinch of salt, but the independent number alone puts Inkling in the same conversation as the better Chinese open models rather than a tier below them.

There is no price, exactly

Here is the wrinkle that a normal launch post glosses over. Thinking Machines does not sell Inkling tokens. There is no api.thinkingmachines endpoint where you drop in a key and get billed per million. The launch post hands you a list of ecosystem partners and a Hugging Face download, and leaves the serving to them. That is a defensible open-source choice, but it means the answer to "what does Inkling cost" depends entirely on which partner you can actually buy from, and most of them are not selling tokens either.

The number that has spread fastest, $1.87 input and $4.68 output, comes from Artificial Analysis' price column, which in turn pulls from Thinking Machines' Tinker platform. Tinker is a fine-tuning and training product with prefill, sample, and train meters. Its sampling rate is not a chat-completions API price, and AA's own footnote admits it is a stand-in used because no first-party API exists. If you quote $1.87/$4.68 as "Inkling's API price," you are quoting the cost of sampling from a fine-tune job, which is a different thing with a different bill.

Strip out the fine-tuning meter and the rent-the-GPU listings, and one real pay-per-token price is left standing.

WhereHow you payPer-token rate
OpenRouter (one provider)Per token, chat API$1.00 / $4.05
Tinker (Thinking Machines)Fine-tune sampling meter$1.87 / $4.68
Fireworks AIOn-demand GPU rentalNo token rate
Baseten, ModalPer GPU-minuteNo token rate
Together AI, DatabricksListed partnerNo token rate

Per million tokens, input / output, as of July 20, 2026. OpenRouter's page notes a single hosting provider behind its listing. The Tinker rate reflects the 64K-context tier with the current 50% launch discount applied. Fireworks, Baseten, and Modal bill for hardware time, not tokens.

One provider is a thin market. It means no price competition to pull the rate down, no redundancy if that provider throttles or pulls the model, and a listed price that could move the moment a second host shows up or the first one reprices. Treat OpenRouter's $1/$4.05 as a real but fragile data point, not a settled rate card.

What a real month costs, if you take that one price at face value

Run the one Inkling price we have against the open and closed models people actually weigh it against. The workload below is a reasoning-leaning month: 15M input tokens, 5M output, no caching, list rates. It is the shape of a small team pushing an agent or a research assistant, not a high-volume classification job.

ModelInputOutput15M in / 5M out
DeepSeek V4 Pro$0.44$0.87$11
Inkling (OpenRouter)$1.00$4.05$35
GLM-5.2$1.40$4.40$43
Claude Sonnet 5 (intro)$2.00$10.00$80
GPT-5.6 Terra$2.50$15.00$113
Kimi K3$3.00$15.00$120

Per million tokens, list rates, rounded to the dollar. DeepSeek V4 Pro is $0.435/$0.87. Claude Sonnet 5's intro rate holds through August 31, then rises to $3/$15. Cross-model per-token comparisons are approximate because tokenizers differ between vendors.

Read that middle of the pack correctly. At $35, Inkling undercuts GLM-5.2, both GPT-5.6 Terra and Kimi K3 by a wide margin, and even Sonnet 5's discounted launch rate. It is not a bargain, though. DeepSeek V4 Pro does the same month for eleven dollars, a third of Inkling's bill, and it is not obviously three times worse at the work. What you are paying the Inkling premium for is a newer architecture, a US-based lab, and a permissive license you can actually build a product on. Whether that is worth 3x DeepSeek depends on your risk tolerance more than your benchmark spread.

The self-host mirage

"Open weights" makes people reach for the self-host calculator, so let us do the math and temper the fantasy. Even though only 41B parameters fire per token, all 975B have to live in GPU memory. The BF16 checkpoint is around 1.9 TB and wants roughly sixteen H200s or eight B300s. The 4-bit NVFP4 checkpoint drops that near 600 GB, which fits on about eight H200s. There is no laptop version of this model.

Turn that into a token cost and the numbers get slippery, so treat what follows as an estimate with the assumptions on the table. Eight H200s rent for somewhere around $20 to $28 an hour on the open cloud market. If you can hold that cluster at high, steady output throughput, the arithmetic lands somewhere near $1.30 to $3.30 per million output tokens. That is a real discount against OpenRouter's $4.05, but only if your traffic is heavy and constant enough to keep eight expensive GPUs busy. The moment your load is bursty, idle GPU time eats the saving and the hosted per-token rate wins. Open weights buy you leverage and control, not a free lunch.

Worth building on yet?

If you want the capability today and you are fine routing through one provider, OpenRouter's $1/$4.05 is a fair price for what Inkling scores, and you can wire it up this afternoon. If you are planning production spend, wait a beat. A single-host market with an unsettled rate and a headline price that keeps getting confused with a fine-tuning meter is not a foundation to budget a quarter on. Give it two or three weeks for Together, Baseten, or Fireworks to post real serverless token rates, and the price will both settle and, almost certainly, fall.

And if the job is routine rather than frontier, none of this is your decision to sweat. DeepSeek V4 Pro sits a third of the price with plenty of capability for extraction, classification, and summarization, and it has a mature multi-provider market behind it. Inkling is the interesting new option for teams that specifically want an open, US-built frontier-adjacent model and can live with a market that is still one provider deep. Drop your own token mix into a cost calculator before you commit, because the only bill that matters is the one your workload produces.

Sources