Kimi K3's weights went up for free on Monday and every host still matches Moonshot's $3/$15 on demand. That is what happens when the smallest machine that boots a model is eight B300s.
Open weights usually start a price war. A lab ships a checkpoint, the commodity hosts race each other to the floor, and within a week the API is the expensive option. Kimi K3 broke that pattern on July 27. Ten providers now serve it, not one undercuts Moonshot's $3.00 in and $15.00 out per million tokens on synchronous requests, and one of them is actually more expensive once you account for caching. The reason is sitting in the download: 1.56 terabytes of MXFP4 weights that will not fit on any Hopper node and not even on eight B200s. Below is the arithmetic on the checkpoint size, the exact machines that can hold it, what those rent for, the sustained throughput where owning beats renting, the one real discount available anywhere, and the clause Moonshot added to this license that was not in K2's.

Photo by Domaintechnik on Unsplash
The release in five lines
- Full checkpoint, ungated, on Hugging Face, timestamped July 27 at 16:29 UTC
- 2.78T parameters, 104B active, 16 of 896 experts plus two shared, 93 layers split 69 KDA and 24 gated MLA
- MXFP4 weights with quantization-aware training from the supervised stage onward, and a 1,048,576-token window
- API live since July 16 at $3.00 input, $0.30 cache hit, $15.00 output, flat across the whole context
- vLLM and SGLang support landed the same day the weights did
A real, complete, day-zero open-weight release of a frontier-class model. Also, for almost everyone reading this, unrunnable.
Ten hosts, one price, and one discount nobody is advertising
Start with the thing that should not be true. We pulled every provider serving K3 and on synchronous requests they all land on Moonshot's own rate. Nobody is using open weights to undercut the lab that released them, and Nebius manages to come out dearer.
| Host | Input / 1M | Cache read | Output / 1M | Measured output speed |
|---|---|---|---|---|
| Moonshot | $3.00 | $0.30 | $15.00 | 32.0 tok/s |
| Fireworks | $3.00 | $0.30 | $15.00 | 164.4 tok/s |
| Together | $3.00 | $0.30 | $15.00 | 55.6 tok/s |
| Baseten | $3.00 | $0.30 | $15.00 | Not published |
| SiliconFlow | $3.00 | $0.30 | $15.00 | Not published |
| DigitalOcean Gradient | $3.00 | $0.60 | $15.00 | Not published |
| Novita | $3.00 | Not published | $15.00 | Not published |
| Atlas Cloud | $3.00 | $0.30 | $15.00 | Not published |
| Nebius | $3.00 | None | $15.00 | 128.1 tok/s |
| Fireworks Priority | $3.75 | $0.375 | $18.75 | Not published |
| Fireworks K3 Fast | $4.50 | $0.45 | $22.50 | Not published |
| Fireworks batch | $1.50 | $0.15 | $7.50 | Asynchronous |
Three things jump out. The first is that Fireworks serves the same model at the same price five times faster than Moonshot does, 164.4 tokens per second against 32.0, and waits 13.3 seconds for the first token where Moonshot takes 160.8. Same $15 a million, roughly a twelfth of the wait, so there is little reason to sit on the first-party endpoint unless you specifically want it. The second is Nebius, which matches the sticker at $3 and $15 but offers no cache-read discount at all. On Artificial Analysis's 7:2:1 traffic mix that works out to $4.20 per million against Moonshot's $2.31, so the one provider that deviates meaningfully deviates upward. DigitalOcean is a smaller version of the same: its docs list prompt caching at $0.60 where everyone else charges $0.30, though the page does not separate cache reads from writes, so read that row as an inference rather than a quote.
The third is the only genuine discount on the board, and it is easy to miss because it is not a K3 price at all. Fireworks bills batch inference at half its serverless rate across every model it hosts, which lands K3 at $1.50 in and $7.50 out. Moonshot publishes no batch tier, and Together excludes Kimi models from its batch API entirely, so this is the single way to pay less than the lab. It is a platform-wide modifier rather than anyone competing on this model, and you pay for it in latency, but if your workload tolerates asynchronous processing it halves the bill.
Then there is who is not here. DeepInfra and Parasail both top out at K2.7-Code, and Groq lists no Moonshot model whatsoever. These are exactly the hosts that normally take an open-weight release and halve the lab's price within days. Novita, which is the same kind of operator, does carry K3, and carries it at $3 and $15 with no discount at all, which if anything makes the point more sharply than the absences do. The commodity end of this market either could not take the model or could not find room under the price. The next two sections are why.
The download is 1.56 terabytes, and the number you have read is wrong
A figure of 594 GB has been repeated across a lot of the quick-turnaround coverage of this release. It cannot be right, and you can see that without checking anything: 2.8 trillion parameters at four bits each is 1.4 TB before you store a single scale factor. So we reconstructed the real size from Hugging Face's own parameter-dtype metadata for the repository, which reports 2,722,740,830,208 MXFP4 elements, 57,179,884,544 BF16 parameters and 11,122,432 in F32.
| Component | How it is stored | Size |
|---|---|---|
| MXFP4 weights | 2.7227T elements at 0.5 bytes | 1,361.4 GB |
| MXFP4 scales | One FP8 per 32 elements | 85.1 GB |
| BF16 tensors | 57.18B parameters at 2 bytes | 114.4 GB |
| F32 tensors | 11.1M parameters at 4 bytes | 0.04 GB |
| Total | 96 safetensors shards | 1,560.9 GB |
Hugging Face independently reports 1,561,018,243,668 bytes of storage for the repo, so the reconstruction lands within 158 MB on 1.56 TB, a hair over 0.01%. Two honest notes on method: the metadata labels that first dtype U8 rather than MXFP4, and the 85.1 GB scales line is our own inference from the OCP microscaling spec, which pairs one FP8 scale with every block of 32. The fact that the total then matches HF's storage figure is what makes us confident both readings are right, and it also settles that the U8 count is unpacked 4-bit elements rather than bytes, since bytes would put the repo at 2.7 TB. While we were in that metadata we also checked two other claims doing the rounds. The repository is not gated, despite several posts saying otherwise. And the architecture is supported in mainstream tooling from day zero, not unsupported: vLLM published a K3 recipe on July 27 and SGLang lists it in the cookbook. Three of the four things most widely repeated about this release are wrong, and they are wrong in a direction that makes self-hosting sound far more approachable than it is.
The smallest machine that boots it
The vLLM team's day-zero writeup states the requirement plainly: at least one 8x B300 node or a GB300 NVL72, with 16x B200 also supported. That is the NVIDIA answer, and it is not the only one: SGLang also lists MI350X and MI355X, and 8x MI355X comes to the same 2,304 GB as 8x B300 with native FP4 on the die. AMD is a real option here that almost nobody is discussing. Line the 1,561 GB checkpoint up against the usual rental configurations and you can see exactly where that rule comes from. This is weights only, before any KV cache, activation memory or runtime overhead, so treat every positive number in the last column as generous.
| Configuration | Total VRAM | Against 1,561 GB |
|---|---|---|
| 8x H100 80GB | 640 GB | Short by 921 GB |
| 8x H200 141GB | 1,128 GB | Short by 433 GB |
| 16x H100 80GB | 1,280 GB | Short by 281 GB |
| 8x B200 180GB | 1,440 GB | Short by 121 GB |
| 16x H200 141GB | 2,256 GB | Fits, 695 GB spare |
| 8x B300 288GB | 2,304 GB | Fits, 743 GB spare |
| 16x B200 180GB | 2,880 GB | Fits, 1,319 GB spare |
The 8x B200 row is the one that stings. It misses by 121 GB, roughly 8% of the checkpoint, which is why vLLM specifies sixteen of them rather than eight. SGLang does list H200 and H100 as supported parts, which reads like an escape hatch until you work out that you need 16x H200 or 24x H100 to physically hold the thing, and Hopper has no native MXFP4 anyway, so those builds pay a dequantization tax on every forward pass. There is no clever quantization trick waiting here either. The checkpoint is already 4-bit; going lower means retraining, and quantization-aware training from the supervised stage is what makes this one work at four bits at all.
What eight B300s cost, and the counterintuitive bit
We checked on-demand published rates across every provider we could reach on July 28. Most of the market has moved to contact-sales for Blackwell Ultra, and only three still print a B300 rate at all. Monthly figures below are 720 hours, so a machine that never sleeps for thirty days.
| Provider and config | Per GPU-hour | Whole machine | 30 days, 24/7 |
|---|---|---|---|
| RunPod, 8x B300 | $7.39 | $59.12/hr | $42,566 |
| Verda, 8x B300 | $7.50 | $60.00/hr | $43,200 |
| Nebius, 8x B300 | $7.85 | $62.80/hr | $45,216 |
| RunPod, 16x B200 | $5.89 | $94.24/hr | $67,853 |
| Nebius, 16x B200 | $7.15 | $114.40/hr | $82,368 |
| CoreWeave, 16x B200 | $8.60 | $137.60/hr | $99,072 |
Notice that the newer, more expensive chip is the cheaper way in. B300 costs more per GPU-hour than B200 at every provider selling both, and you need half as many of them, so the 8x B300 node beats 16x B200 by $31 to $78 an hour. If you have been assuming Blackwell Ultra is the option you skip on price, this workload inverts that. One caveat: RunPod lists B300 at 288 GB, Verda lists theirs at 268 GB, and 8x 268 GB is 2,144 GB, which still fits but leaves 583 GB rather than 743 GB for KV cache. Check the SKU before you assume the headroom.
Full GB300 NVL72 racks, which is what Baseten is actually running in production at nine replicas per rack, are effectively unpriced in public. Every provider we checked routes them to sales. The only published rate we found was Verda at $8.62 per GPU-hour.
The load where owning beats renting
Take the cheapest verified machine, RunPod's 8x B300 at $59.12 an hour. Rented continuously that is $42,566 a month, and it is a fixed cost whether you push a billion tokens through it or none. So the only question that matters is how many tokens you can keep it busy with. Divide the hourly rate by throughput and you get a real cost per million output tokens to set against the API's $15.
| Sustained output throughput | Where the figure comes from | Your cost / 1M out |
|---|---|---|
| 111 tok/s | vLLM batch 1, TP8, one user | $147.95 |
| 331 tok/s | Batch 1 with DSpark speculative decoding | $49.61 |
| 1,095 tok/s | Break-even against the $15 API rate | $15.00 |
| 7,109 tok/s | Break-even against a $2.31 blended bill | $2.31 |
| 16,000 tok/s | vLLM ceiling, 2K per GPU per second | $1.03 |
Two break-even rows because there are two honest ways to state the comparison. Against the $15 output rate alone you need 1,095 output tokens per second sustained, which is 137 per GPU. That is a low bar next to a 2,000-per-GPU ceiling, and it is why the answer here is not the flat "self-hosting always loses" you get with smaller open-weight models. But output tokens are not the whole bill. Artificial Analysis puts K3's blended rate at $2.31 per million on a 7:2:1 mix of cache hits, fresh input and output, and we get the same figure recomputing it from the rate card. Set your machine against $2.31 rather than $15 and break-even jumps to 7,109 tokens per second, or 889 per GPU, held around the clock for thirty days.
Which of those you should use depends on how cacheable your traffic is, and most real agent workloads are very cacheable, so most readers should be looking at the harder number. A caveat that applies to every throughput row above, not just the last one: vLLM measured all of these on GB300 NVL72, and we are pricing them against a rented 8x B300 node because that is the cheapest thing you can actually buy by the hour. The two are not the same machine. Tokens per GPU per second may also include prefill, so treat 2,000 as a ceiling you approach rather than a rate you get. The rough shape: if you are running fewer than a few thousand output tokens per second every second, the API is cheaper and it is not close.
K3 writes long, and that eats most of its price advantage
Worth pausing on this before anyone picks K3 on the rate card alone. Artificial Analysis publishes cost per task on its Intelligence Index, which is the number the per-million rate is supposed to approximate and frequently does not.
| Model | Blended / 1M | Cost per Index task |
|---|---|---|
| DeepSeek V4 Pro | $0.18 | $0.04 |
| GLM-5.2 | $0.90 | $0.32 |
| Kimi K3 | $2.31 | $0.94 |
| GPT-5.6 Sol | $4.35 | $1.04 |
| Claude Opus 4.8 | $3.85 | $1.80 |
Read the two columns against each other and they do not track. Opus 4.8 blends cheaper than Sol and costs 73% more per task. K3 is the sharpest case: its tokens are 1.88 times cheaper than Sol's on a blended basis, and its tasks are 9.6% cheaper. The gap between those two numbers is verbosity. Running the Index cost $2,437.41 and generated 130 million tokens against a 99 million median across the field, which Artificial Analysis calls very verbose. K3 scores 57 on the Index, which is genuinely close to the frontier, and it talks its way there. Worth noting in Moonshot's defence that this is an improvement: K3 spent about 21% fewer output tokens on the same suite than K2.6 did.
The practical version. On the rate card K3 looks like it should cost you 53% of what Sol costs. On completed tasks you pay 90%. Roughly a fifth of the saving the sticker implies actually survives contact with the workload, so if you are switching to K3 to cut a bill, model it on your own traces. The cost calculator will take your real token counts if you have them.
The license has two dollar figures in it
"Modified MIT" is a fair shorthand. The file opens with verbatim MIT grant text and then adds five clauses, two of which are thresholds rather than restrictions. Clause 2 says that if you operate the model as a service and your aggregate revenue passes $20 million over any consecutive twelve months, you need a separate agreement with Moonshot. Clause 3 says that above 100 million monthly active users or $20 million in monthly revenue, you have to show "Kimi K3" prominently in your interface. Clause 4 exempts internal use entirely, along with access through Moonshot's own products or its certified inference partners.
For anyone self-hosting for their own product rather than reselling inference, none of it binds. It is aimed at hyperscalers reselling the model, not at you. Worth being precise about one thing, though, because the "same as K2" framing is going around and it is not right: K2's license carries the attribution paragraph and nothing else. It has no model-as-a-service clause at all. The $20M aggregate-revenue MaaS trigger is new in K3, so this license is meaningfully tighter than its predecessor, not a copy of it.
That said, the practical reading does not change. The license is not what stops you self-hosting K3. The $42,566 a month is.
Insurance, not savings
Not for saving money, for most people. If your output volume is below roughly a thousand tokens per second sustained, and for the overwhelming majority of teams it is, renting eight B300s to serve yourself costs more than buying the same tokens from Fireworks at five times Moonshot's speed for the same price. The weights are worth having for data isolation, air-gapped deployment, fine-tuning, and as insurance against an endpoint whose terms or availability can change. Moonshot suspended new consumer subscriptions on July 19, three days after launch, when GPU capacity ran out, which is a useful reminder that a published price you cannot buy at is not really a price.
What I keep turning over is what this does to the meaning of an open-weight release. For two years the pattern has been reliable: weights drop, DeepInfra and Novita and friends undercut the lab within days, and the API becomes the convenience tax. K3 is the first release big enough that the pattern breaks. The weights are free and the price does not move, because the barrier stopped being the license and became the machine. Publishing a checkpoint that needs $42,000 a month of Blackwell Ultra to run is open in the letter and closed in the effect, and I do not think that is a criticism of Moonshot so much as a description of where model sizes have got to.
If you want the cheap end of this market, it is not here. GLM-5.2 blends to $0.90 and DeepSeek V4 Pro to $0.18, and both run on hardware you can actually rent by the hour. K3 is a frontier model with a frontier machine attached, and the honest version of the self-hosting question is not "can I save money" but "do I have 8x B300 of steady demand." Almost nobody does.
Put your own numbers in
Kimi K3 sits in the TokenCost catalogue next to the rest of the frontier. Compare it against whatever you are running now with your real token volumes.
Sources
- Hugging Face, moonshotai/Kimi-K3 for the model card, architecture and the July 27 16:29 UTC timestamp. The API metadata endpoint supplies the parameter-dtype counts the 1,560.9 GB reconstruction is built from, and the gated: false flag.
- The Kimi K3 license text for the MIT grant and the five added clauses, including the $20M revenue and 100M MAU thresholds, read against K2's license , which has no numbered clauses and no model-as-a-service provision.
- Moonshot pricing docs for the $3.00 / $0.30 / $15.00 first-party rate card and the absence of any long-context, batch or off-peak tier.
- vLLM, Kimi K3 day-zero support for the 8x B300 minimum, the 16x B200 alternative, and the batch-1 and speculative-decoding throughput figures. The SGLang cookbook for its supported-hardware list, including the AMD MI350X and MI355X parts. The 2K-per-GPU-per-second ceiling is vLLM's figure, not SGLang's; the SGLang cookbook publishes no throughput numbers.
- Baseten for the production topology of 8x GB300 across two nodes at nine replicas per NVL72 rack.
- RunPod, Verda, Nebius and CoreWeave for on-demand B300 and B200 rates, all checked July 28, 2026. Together, Lambda and Crusoe publish no B300 rate at all.
- Fireworks, Together, Baseten and OpenRouter for the third-party rate cards and measured throughput. Fireworks' batch pricing is the "50% of serverless pricing on both input and output" line in its serverless docs; Together's batch documentation excludes Kimi models. Novita's listing comes from its own models API rather than its catalogue pages, which had not caught up. DigitalOcean's $0.60 comes from its own docs, not the OpenRouter relay, and is listed there as prompt caching without separating reads from writes.
- Artificial Analysis for the Intelligence Index score of 57, the $2.31 blended rate, the cost-per-task figures, and the 130M versus 63M median token counts.
- Caixin for the July 19 suspension of new Kimi signups after GPU capacity ran out.
- Anthropic, OpenAI, Z.ai and DeepSeek for the comparison rate cards. Blended figures are our own calculation on Artificial Analysis's 7:2:1 cache-hit, input, output weighting.