Skip to main content
TokenCost logoTokenCost
Model ReleaseAugust 29, 2026·12 min read

Tencent priced its new flagship at 6 yuan per million input tokens and printed $0.834 next to it. That conversion uses 7.196 to the dollar, the market closed Friday at 6.73, and the gap is why every catalogue now says Tencent raised its price 6.32x when it raised it exactly 6.

Hy4 preview arrived on Thursday morning with 770 billion parameters, 49 billion of them active, a window just over a million tokens and an Apache 2.0 licence. The number everyone copied was $0.834. It is a real price and you can pay it, but it is not a price anybody chose. It is six yuan, divided by an exchange rate that was already stale when the page went up, and once you know that, the comparison everyone is making between this model and the one it replaces is off by five percent in Tencent's favour.

Red LED BEST RATES currency exchange sign glowing in a dark shop window at night

Photo by Jon Cellier on Unsplash

One rate card, published twice by the same company on the same day

Input¥6.00divided by 7.196$0.834
Output¥18.00divided by 7.196$2.501
Cache hit¥0.30divided by 7.196$0.042

Per million tokens. The left column is Tencent Cloud's price table and Tencent's Chinese announcement. The right column is Tencent's English announcement, and it is what OpenRouter, every tracker and every writeup published on Thursday. Friday's close on the dollar against the yuan was about 6.73.

The 7.196 is not printed anywhere. We recovered it by asking which single rate turns all three yuan figures into all three dollar figures once each is rounded for display, which lands in a window between 7.1957 and 7.1986, or 0.040% wide. Spot rate from Investing.com's August 28 close of 6.7267, corroborated within 0.02 by Trading Economics, Wise and OFX.

Six yuan, eighteen yuan, and thirty fen

Tencent listed Hy4 preview on OpenRouter at 06:09 UTC on August 28 and put out an announcement in two languages the same day. The Chinese one gives the price as 6 yuan per million input tokens, 18 yuan output and a cache hit from 0.3 yuan. The English one gives it as $0.834, $2.501 and $0.042. Tencent Cloud's TokenHub price table, which is the document you would actually be billed against, carries the row as 6, 18 and 0.3 in the Guangzhou region.

So there is no detective work in the first half of this. Tencent published both currencies itself, and OpenRouter copied the dollar column rather than converting anything. What nobody did was note which of the two is the price and which is the derivative. Every English writeup we found, including TechNode's, which is the most careful of them, quotes $0.834 and $2.501 and stops there.

That matters because a converted price behaves differently from a set one. It moves when nothing about the product moves. It cannot be compared cleanly against another converted price unless both used the same rate on the same day. And it carries a rounding error that you can measure, which turns out to be the most useful thing on the page.

The rounding residue is the evidence

Look at what the dollar card says about itself. Output divided by input is 2.501 over 0.834, which is 2.998801. Cache divided by input is 0.042 over 0.834, which is 5.0360%. Both are close to a round number and neither is one. In yuan they are exactly 3.000000 and exactly 5.0000%.

Nobody sets a rate card where output is 2.9988 times input. What produces a figure like that is three separate numbers each divided by the same rate and each rounded to three decimals on the way out. The residue is small, 0.04% on the output ratio and 0.72% on the cache ratio, and it is the fingerprint. It also does the real work here, because those three inexact ratios constrain the rate from three directions at once. Input alone allows anything from 7.1899 to 7.1986. Add output and the window closes to 7.1957. Cache is consistent with both and too coarse to narrow it further. What survives is 0.040% wide, which is a tighter reading of a number Tencent never printed than we normally get on numbers vendors do print.

We ran the same test against all 396 models in OpenRouter's catalogue, and the honest result is that it usually fails to be decisive. Most cards are exactly proportional, so any round card at some rate fits and the rate is unrecoverable. Hy3 is the case in point, and it is the next section but one.

Friday's close was 6.73, and the difference is yours to keep

A frozen rate is only interesting if it has drifted, and this one has. The dollar closed against the yuan at about 6.73 on August 28, the day Hy4 preview shipped. Tencent's conversion used 7.196. That is 7.0% above the market on the day the page went up, and above OFX's six-month average of 6.804 as well, so this is not a rate that was current last week and slipped.

Run it the other way and the drift is a discount. Six yuan converted at Friday's close is $0.892. Tencent charges $0.834. Buying Hy4 preview in dollars costs 6.49% less than buying the identical tokens in yuan, and the same holds on output at 6.53% and on cache reads at 5.82%. If you are an international customer, the stale rate is working for you. On a $200 monthly agent bill it is worth about thirteen dollars, which is not much; the reason to know about it is the direction it points.

Because the mechanism runs both ways. Tencent can refresh the conversion at any time, and doing so would take input from $0.834 to $0.892, a rise of 6.95%, with no announcement, no changelog entry and no change to the yuan card. We have written up eight API price increases with published expiry dates this month. This is a ninth kind, and it is the only one with no date attached to it at all.

Hy3 is frozen at a different rate, and that is the entire 6.32

Here is where the arithmetic stops being trivia. Hy3 sells for 1 yuan input, 4 yuan output and 0.25 yuan on a cache hit. Its dollar card is $0.132, $0.528 and $0.033. Divide those into each other and the implied rate is somewhere between 7.5686 and 7.5829, which is nowhere near Hy4 preview's 7.196 and further still from the market. Two models from one company, sitting one row apart in the same price table, converted at rates 5.3% apart.

So the headline comparison is contaminated. Every catalogue that carries both models says Tencent went from $0.132 to $0.834, which is 6.3182x. Tencent went from 1 yuan to 6 yuan, which is 6.0000x. The 0.3182 in between is a currency market, and it is bigger than several of the price moves we have covered as news this year.

LineHy3 to Hy4, in yuanYuan multipleSame move, in dollarsDollar multiple
Input¥1 to ¥66.0000x$0.132 to $0.8346.3182x
Output¥4 to ¥184.5000x$0.528 to $2.5014.7367x
Cache hit¥0.25 to ¥0.301.2000x$0.033 to $0.0421.2727x

Every dollar multiple overstates the yuan one, by 5.30% on input, 5.26% on output and 6.06% on the cache line. They are not identical because each of the six dollar figures rounded independently, and the smallest number rounds worst.

One caveat, and we would rather state it than bury it: the test that pinned Hy4 preview's rate to a 0.040% window pins Hy3's only to 0.190%, because Hy3's card is exactly proportional. Its output is precisely four times its input and its cache is precisely a quarter, in both currencies, so the ratios tell you nothing and only the display precision does. We can bracket Hy3's frozen rate. We cannot say which day it was taken from, and we could not find Tencent stating that anywhere.

One more caveat on the sourcing, since it is the kind of thing that can quietly invalidate a comparison. Tencent's yuan table is the Guangzhou region and its dollar table is Singapore, so these are not formally the same SKU. The two announcement pages carry both currencies with no region attached at all, which is what most readers will act on, and it is the pair we are comparing.

Tencent is not the odd one out, which is the actual problem

We went looking for how common this is, expecting to find that Chinese vendors price in yuan and Western catalogues convert. That turns out to be mostly wrong, and the truth is stranger: the vendors do the converting themselves, each at its own fixed rate, and those rates are all over the place. Here is every one we could read straight off a first-party page, next to a market that closed the week at 6.727.

VendorWhat it coversRate in useAgainst 6.727
TencentHy37.57612.6% above
AlibabaQwen Max family, Beijing card7.2718.1% above
TencentHy4 preview7.1946.9% above
AlibabaQwen Flash, Beijing card7.0685.1% above
MiniMaxEvery model, every tier7.0004.1% above
DeepSeekEvery model, both windows6.8181.4% above

Every rate in that table is exact to five significant figures across every line item on the relevant page, which is what tells you it is a conversion rather than two lists that happen to be near each other. MiniMax is the purest case: a flat 7.000 on every model, every tier, cache reads and cache writes, before and after its permanent half-price discount, down to a cache-write line that is precisely $0.375 times seven. DeepSeek runs 6.818 uniformly and is the only vendor here within a couple of percent of the market. Alibaba is the most tangled, publishing genuinely independent yuan and dollar price lists for its Beijing and Singapore regions and then rendering each one in the other currency at a different fixed rate, four of which are live on its documentation right now.

Two vendors break the pattern, and they are worth naming because they show it is a choice. Z.ai lists GLM-5.3 at ¥8 and $1.40, which implies 5.71 on input, 6.36 on output and 7.69 on the cache line. No single rate fits, because those are two separately authored price lists, each rounded to numbers a human would pick in its own currency. Moonshot does the same. The consequence shows up in the marketing: Z.ai's Chinese documentation says Flash is priced at a tenth of GLM-5.3, which is true at ¥0.8 against ¥8 and false at $0.15 against $1.40.

The cache line looks flat and is the biggest change on the card

In dollars, cached input went from $0.033 to $0.042. That reads as the boring row, up 1.27x while input went up 6.32x, and it is easy to skim past. Read it as a fraction of the thing it discounts and it is the most consequential edit Tencent made.

On Hy3 a cache hit costs a quarter of a fresh input token. On Hy4 preview it costs a twentieth. Tencent widened its cache discount from 4x to 20x while raising everything else, which is a deliberate shape: it makes the model much cheaper for the workload it is being sold for, agents that resend the same context every turn, and leaves it expensive for one-shot work that never repeats. Anthropic, Google and Moonshot all sit at 10x on the same line. DeepSeek V4 Pro is at 30x and Meituan's LongCat-2.0 at 50x, so Hy4 preview is now in the aggressive half of the field rather than the bottom of it.

The dollar card, incidentally, says 19.86x rather than 20x. That is the rounding again, and it is a good demonstration of why the currency question is not academic: a reader working from the English page would conclude Tencent set the cache rate at 5.04% of input, which is a strange number to choose, instead of 5%, which is obviously the number somebody chose.

An agent that re-reads one repository, priced eleven ways

Tencent is selling this for coding agents and tool-use workflows, so price it as one. Take 20,000 requests a month where each turn resends 60,000 cached tokens of repository context, adds 3,000 fresh input tokens and produces 2,000 output tokens. That is 1.2 billion cached, 60 million fresh and 40 million output over a month, and it is the shape that Hy4 preview's new cache rate is built for.

ModelMonthly billNote
Hy3, 16:00-00:00 UTC$42.90Tencent's evening card, on OpenRouter only
Hy3$68.64262,144-token window
LongCat-2.0$73.20Cache reads at 2% of input
MiniMax M3$138.00
DeepSeek V4 Pro 0813$145.20Off-peak card; doubles at peak
Hy4 preview$200.48¥1,440 on Tencent's own card
Gemini 3.7 Flash$285.00Doubles January 1
GLM-5.3$572.00Named in Tencent's own comparison
Qwen3.8-Max$660.00
Kimi K3$1,140.00Named in Tencent's own comparison
Claude Opus 5$1,900.00

Hy4 preview lands at $200.48, which is roughly a third of GLM-5.3 and about a sixth of Kimi K3, the two models Tencent chose to name in its own evaluation. Against Opus 5 it is close to a tenth. It is also 2.92 times the bill on Hy3, and in yuan it is 2.77 times, which is the same 5% wedge showing up in the totals rather than the rate card.

The row worth arguing with is DeepSeek V4 Pro at $145.20. That is its off-peak card, and it doubles for seven hours a day, so an agent that runs on a schedule you do not control will land somewhere between $145 and $290 while Hy4 preview stays at $200.48 around the clock. Flat pricing is worth something on a bill you cannot time.

What the million-token window is actually for

Change the workload to something with no cache and a very long prompt, a thousand passes over a 400,000-token codebase producing 8,000 tokens each, and the ordering rearranges. Hy4 preview costs $353.61. Gemini 3.7 Flash costs $330.00 and DeepSeek V4 Pro costs $279.84, so on raw long-context reading Tencent is the third cheapest of the six models that can hold the prompt at all.

Model1,000 passesContext window
DeepSeek V4 Pro 0813$279.841,048,576
Gemini 3.7 Flash$330.001,048,576
Hy4 preview$353.611,048,576
GLM-5.3$595.201,310,720
Qwen3.8-Max$848.001,000,000
Kimi K3$1,320.001,048,576
Hy3Cannot run it262,144

The last row is the point of the whole exercise. Hy3 does not appear with a number because it cannot take the prompt: 262,144 tokens against 400,000. Whatever you think of the price move, the window went up 4x and that is the part you cannot get by paying Hy3 more. Output is capped at 64,000 tokens, which is the lowest ceiling in the table and worth checking against your longest generation before you migrate anything.

The evening discount did not follow the model across

Ask OpenRouter's endpoint API for Hy3 and the Tencent listing comes back with a time-based override: between 16:00 and 00:00 UTC it bills $0.0825, $0.33 and $0.020625, which is 0.625 times the standard card on every line. That is a 37.5% cut for eight hours a day, and it is why the top row of the agent table is $42.90 rather than $68.64. The same call for Hy4 preview returns no override at all. It is flat, one seller, one card, all day.

We went looking for that window on Tencent's side and could not find it. Tencent does document a peak and off-peak scheme, but it is written for the DeepSeek models it resells rather than its own, its peak hours are Beijing 09:00 to 12:00 and 14:00 to 18:00, and its price table shows Hy3 and Hy4 preview both flat. OpenRouter's Hy3 window converts to Beijing midnight through 08:00, which sits inside Tencent's off-peak definition without matching its edges. So treat the $42.90 as a real price you can pay through OpenRouter and an unexplained one, and do not assume a similar discount is coming to Hy4 preview.

One seller, and a file that says it will not stay that way

Hy4 preview has exactly one endpoint on OpenRouter today, Tencent itself, serving fp8 at 100% uptime. Hy3 has seven, running from GMICloud at $0.126 up to AtlasCloud at $0.20, a spread of 1.59x on identical weights. That is the normal life cycle of an open-weights release and Hy4 preview is one day into it.

The weights are on Hugging Face, ModelScope, GitCode and CNB in BF16 and FP8, under Apache 2.0 for both code and weights with no separate commercial agreement. That is a meaningfully more permissive licence than Hy3 preview shipped under, which was Tencent's own community licence, and it is what makes the single-seller situation temporary. When other hosts arrive they will price in dollars because their costs are in dollars, and the yuan question quietly stops applying to anyone who is not buying from Tencent.

Scores, and who is holding the ruler

The number Tencent led with is a blind evaluation it ran itself: 163 internal experts scoring 203 engineering tasks, where Hy4 preview averaged 2.99 out of 4.00 against GLM-5.3 at 2.92 and Kimi K3 at 2.94. The head-to-head splits are more informative than the averages. Against GLM-5.3 it won 46.8%, tied 12.8% and lost 40.4%. Against Kimi K3 it won 51.2%, tied 7.9% and lost 40.9%. Those are narrow margins from a vendor grading its own launch, and they are the only comparison against those two models that exists right now.

BenchmarkHy4 preview
GPQA Diamond92.3
Terminal-Bench 2.185.4
MCP-Atlas83.7
SWE-bench Multilingual82.9
Toolathlon-Verified74.1
SWE-bench Pro65.7
DeepSWE64.3
Skillsbench V1.162.9
HLE, with tools55.4
APEX-Agents, pass@137.1

These come from Tencent's model card, and the ones we could cross-check appear on DataLearner's page with the same values. The DeepSWE figure is the one that stands out: 64.3 against Hy3's 28.0 on the same benchmark. LMArena has it around fifth in Code Arena WebDev at 1633 points, up from Hy3 at thirty-first, though that score is an early automatic evaluation rather than live human votes and Arena says so on the post. There is no Artificial Analysis entry yet, which means the number this blog usually cares about most, what it costs to finish a fixed suite rather than what it costs per token, does not exist for this model. We will revisit when it does.

Read the column on the left

Hy4 preview is a good card for what it is aimed at. Two hundred dollars a month for an agent workload that costs $572 on GLM-5.3 and $1,140 on Kimi K3, with a window four times its predecessor's and a cache discount five times deeper, is a real offer, and the weights are Apache 2.0 if you would rather not rent it at all.

What we would not do is put $0.834 in a spreadsheet next to $0.132 and draw a conclusion about Tencent's pricing strategy from the ratio. Those two numbers were minted on different days at different rates, and the 5.30% between what they say and what Tencent did is larger than several of the price changes that got their own headlines this year. When a vendor prices in one currency and publishes in two, the published one is a translation, and translations drift. Six, eighteen and thirty fen are the numbers Tencent chose. Everything else on the English page is arithmetic.

Two announcement pages, both sides of five price lists, and a rate we had to bracket

  • Tencent: Hy4 preview announcement, Chinese - The yuan card, stated as 6 yuan input, 18 yuan output and a cache hit from 0.3 yuan per million tokens. Also the source of 770B total and 49B active parameters, the context length over 1M, the two-week free trial on WorkBuddy and CodeBuddy, and the 163-expert, 203-task blind evaluation at 2.99 against GLM-5.3 on 2.92 and Kimi K3 on 2.94
  • Tencent: Hy4 preview announcement, English - The same launch with the dollar card, $0.834 input, $2.501 output and $0.042 cache hits. Published the same day as the page above, which is what makes the pair of them the whole argument of this post rather than an inference from one of them
  • Tencent Cloud: TokenHub model prices - The billing document. Hy4 preview reads 6 / 18 / 0.3 and Hy3 reads 1 / 4 / 0.25 in the Guangzhou region, both flat, with the peak and off-peak columns empty on each. Hy3 preview is on the same table with a length-tiered card, 1.2 / 4 / 0.4 below 16K input, 1.6 / 6.4 / 0.6 to 32K and 2 / 8 / 0.8 above it, and Tencent publishes no dollar figure for it anywhere. The peak scheme documented on Tencent's side covers the DeepSeek models it resells, at Beijing 09:00-12:00 and 14:00-18:00, and from August 29 it stops applying at weekends
  • OpenRouter: Hy4 preview endpoint detail - Pulled August 29, 2026. One endpoint, Tencent, fp8, at $0.834 / $2.501 / $0.042 with a discount field of 0 and no time-based override. Source of the 1,048,576-token window, the 64,000-token output cap, the 06:09 UTC listing timestamp, and the reasoning block reading optional with efforts of high, low and none defaulting to high. The equivalent Hy3 call returns seven sellers and the 16:00-00:00 UTC override at 0.625x
  • Hugging Face: tencent/Hy4-preview - Apache 2.0 on code and weights, BF16 and FP8 variants, 78 layers with 256 routed experts and top-8 routing, and the benchmark table behind GPQA Diamond 92.3, SWE-bench Multilingual 82.9, SWE-bench Pro 65.7, DeepSWE 64.3 and Skillsbench V1.1 62.9
  • DeepSeek, MiniMax, Alibaba and Z.ai: pricing pages in both currencies - The cross-vendor table. DeepSeek publishes the same page at a /zh-cn path and an English one, and every cell divides to 6.8182. MiniMax does the same across platform.minimaxi.com and platform.minimax.io at a flat 7.000. Alibaba runs help.aliyun.com against alibabacloud.com, where the Qwen Max family converts at 7.271 and Qwen Flash at 7.068 on the Beijing card. Z.ai is the counter-example: open.bigmodel.cn prints GLM-5.3 at ¥8 / ¥28 / ¥2 and docs.z.ai prints $1.40 / $4.40 / $0.26, which implies 5.71, 6.36 and 7.69 on the three lines and therefore no conversion at all. Both of Z.ai's pages and Alibaba's Singapore tables render in JavaScript, so those four rows were read from a rendered page rather than raw HTML
  • Investing.com: USD/CNY historical data - The August 28 close of 6.7267, with 6.7199 on the 27th and 6.7229 on the 26th. Trading Economics reads 6.7298 for the 28th, Wise 6.729 and OFX 6.73978 with a six-month average of 6.80419. We used 6.727 throughout and every figure derived from it would move less than a tenth of a percent on any of the four
  • TokenCost: Hunyuan HY3 Preview pricing - Our May 11 post on the previous generation, which flagged at the time that the figure in our own catalogue was an OpenRouter listing rather than a Tencent rate card. This post is the belated answer to that flag, and the catalogue entry has been updated alongside it
  • What we could not establish. The 7.196 is our reconstruction, bracketed to 7.1957-7.1986 by the display precision of three numbers, and Tencent publishes no rate anywhere; if it applies some internal fixing convention rather than a spot quote, the window is right and the story behind it is not ours to tell. We could not date either frozen rate, only bound them. We could not find any Tencent page documenting the Hy3 evening discount that OpenRouter applies, so its origin is unverified and the $42.90 row rests on OpenRouter alone. The benchmark scores are Tencent's own or DataLearner's transcription of them, with no independent run behind any of them, and the full table in Tencent's repository is published as images rather than text, so we could not check AIME, LiveCodeBench, tau-bench or BFCL at all. And there is no Artificial Analysis entry, which means nothing here prices what a finished task costs, only what a token costs