American teams now route a third of their tokens to Chinese models to save money. The plot twist in July's pricing is that China's best model costs as much as the US flagships it beats.
CNBC put a number on the shift this month: since February, US companies have made up more than 30 percent of Chinese-model token traffic on OpenRouter, week after week. The reason is not loyalty or politics. It is the invoice. But the same July that made that headline also gave us Kimi K3 at $3/$15, the priciest model any Chinese lab has ever charged for. So the tidy "Chinese equals cheap" rule of thumb is now half wrong. We priced nine models on one real workload and ranked them by what you actually pay.

Photo by Gábor Szűts on Unsplash
The three-line version, before the tables:
- - The cheapest coding tokens on the market are Chinese and it is not close: DeepSeek V4-Pro finishes a heavy month for $122 while Opus 4.8 wants $2,000.
- - The most expensive Chinese model, Kimi K3, is not cheap at all. It costs more than Grok 4.5 and GPT-5.6 Luna, two US models, and 2.6 times what GLM-5.2 costs.
- - Nationality stopped predicting price. There are US and Chinese options at nearly every point on the curve now, so the choice comes down to your workload, not a flag.
The rate cards, side by side
| Model | Origin | Input / output per 1M | Context |
|---|---|---|---|
| DeepSeek V4-Flash | China | $0.14 / $0.28 | 1M |
| DeepSeek V4-Pro | China | $0.435 / $0.87 | 1M |
| GPT-5.6 Luna | US | $1 / $6 | ~1M |
| GLM-5.2 | China | $1.40 / $4.40 | 1M |
| Grok 4.5 | US | $2 / $6 | 500K |
| Kimi K3 | China | $3 / $15 | 1M |
| Claude Opus 4.8 | US | $5 / $25 | 1M |
| GPT-5.6 Sol | US | $5 / $30 | ~1M |
| Claude Fable 5 | US | $10 / $50 | 1M |
USD per million tokens, standard tiers. DeepSeek rates shown are off-peak; from the mid-July GA launch they double during Beijing peak hours. Grok bills a higher, unpublished rate above 200K tokens. Sources: OpenAI, Anthropic, xAI, DeepSeek, Moonshot, and Z.ai pricing, mirrored on the TokenCost pricing page.
The cheap end of the market is Chinese, full stop
Pick one workload and hold it constant, because per-token rates are easy to misread. Say a coding-agent team that pushes 200M input tokens and 40M output tokens through the API in a month. That is an input-heavy shape, which is what agentic coding looks like once you count repo context, file reads, and tool results. Here is the monthly bill on each model, cheapest first.
| Model | Origin | Monthly cost | vs Opus 4.8 |
|---|---|---|---|
| DeepSeek V4-Flash | China | $39 | 51x cheaper |
| DeepSeek V4-Pro | China | $122 | 16x cheaper |
| GPT-5.6 Luna | US | $440 | 4.5x cheaper |
| GLM-5.2 | China | $456 | 4.4x cheaper |
| Grok 4.5 | US | $640 | 3.1x cheaper |
| Kimi K3 | China | $1,200 | 1.7x cheaper |
| Claude Opus 4.8 | US | $2,000 | baseline |
| GPT-5.6 Sol | US | $2,200 | 1.1x pricier |
| Claude Fable 5 | US | $4,000 | 2x pricier |
The bottom of that table is where the OpenRouter story comes from. DeepSeek V4-Pro finishes the month for $122 while Opus 4.8 charges $2,000 for the same token counts, a 16x gap that shows up directly on a startup's card statement. V4-Flash makes it absurd: $39 against Fable 5's $4,000, over 100 to 1, so a full year on V4-Flash costs less than a single week of Fable at this volume. And GLM-5.2 slots in at $456 with an MIT license, so a team that wants to self-host can skip the API entirely. If your bar is "good enough coding at the lowest price," you have not needed a US model for months.
Kimi K3 is the line where the discount runs out
Now look at row six. Kimi K3 is Chinese, it launched July 16, and it is the best-scoring model on this list that is not a US flagship. It sits fourth on the Artificial Analysis Intelligence Index and first on Arena's frontend code board. Moonshot priced it accordingly: $3 input, $15 output, up 3.2x on input from K2.6's $0.95/$4.00 three months earlier. On our workload that is $1,200 a month.
Read that against its own neighbors and the pattern everyone assumes just breaks. K3 costs 2.6x what GLM-5.2 costs and close to ten times what DeepSeek V4-Pro charges, and both of those are also Chinese. It runs almost double Grok 4.5 and nearly triple GPT-5.6 Luna, both American. The only thing K3 undercuts is the three US flagships above it, and even there the saving over Opus 4.8 is 40 percent, not the 94 percent you get from DeepSeek. Moonshot did not build a cheaper Opus. It built its own Opus and charged Opus money for it.
That matters beyond one model, because it says the price war has a shape. Chinese labs drove the floor down toward zero and kept it there. The ceiling is a different fight, and the first Chinese model to genuinely reach it decided the reward for reaching it was pricing power, not a race to the bottom. The cheap-China assumption held right up until a Chinese model was actually the smartest thing in the room.
The rate card lies in both directions
Two footnotes move real money, and they cut opposite ways. First, DeepSeek's GA launch in mid-July added peak-hour pricing. Every rate doubles during Beijing business hours, 09:00 to 12:00 and 14:00 to 18:00. A V4-Pro month run entirely in that window is roughly $244, not $122. Still eight times cheaper than Opus, but if your batch jobs happen to fire during Chinese afternoons, you are quietly paying the surcharge. Schedule around it and the $122 number holds.
Second, and in Kimi's favor: sticker price overstates what K3 costs to finish a task. Artificial Analysis measured its per-task cost at $0.94 against Opus 4.8's $1.80, because K3 spends fewer output tokens getting there even though each one is expensive. So the $1,200 line above is the worst-case reading. Price it by finished work instead of raw tokens and K3 closes a chunk of the gap to the models above it. The rate card is where you start, never where you stop.
So which side of the border should your tokens live on
If cost is the binding constraint and your work is coding or agents, the answer has been the same for a while and this month did not change it: DeepSeek V4 or GLM-5.2, off-peak, and you pocket the 4x to 50x. The reasons not to are real but they are not about capability. Data residency, a peak-hour surcharge you have to schedule around, and procurement rules that some US shops simply cannot clear. None of those show up in the price column, so decide them before the price tempts you.
If you want the top of the coding charts, the honest read is that you are choosing between Kimi K3 and the US flagships on merits other than nationality, because K3 priced itself into their bracket. And if your traffic is mid-tier production work, the sleeper on this table is GPT-5.6 Luna at $440, a US model that undercuts every Chinese option except the two DeepSeek tiers. The map is no longer US-expensive, China-cheap. It is a spectrum with players from both countries at almost every price point.
All of these are sticker rates on one made-up workload. Your bill turns on your own input-output split, how often each model needs a retry, and whether you can cache. Drop your real numbers into the cost calculator and line any two of these up on the pricing page. The scoreboard tells you where to look. Your own traffic tells you where to land.
Sources
- - CNBC on US usage of Chinese models: cnbc.com
- - Artificial Analysis, Kimi K3 evaluation: artificialanalysis.ai
- - Simon Willison on Kimi K3: simonwillison.net
- - Related: Kimi K3 pricing breakdown
- - Related: DeepSeek V4 peak-hour pricing
- - Related: GLM-5.2 vs GPT-5.5 coding cost
- - TokenCost pricing page: tokencost.app/pricing