Skip to main content
TokenCost logoTokenCost
ResearchJuly 23, 2026·9 min read

Claude Opus 4.8 just slipped off the LLM value frontier. A $3 model scores higher and costs a quarter less to run.

Rank every current model by its Artificial Analysis Intelligence Index and by what Artificial Analysis actually paid to run that eval suite, and the July 2026 board stops matching the price list. Kimi K3 now scores one point above Opus 4.8 and costs $1,000 less to run the benchmark. Grok 4.5 quietly dominates three models at once. And the single most expensive point of quality on the board belongs to Claude Fable 5, which still holds the top score but charges double to earn it.

Abstract blue data-flow visualization representing LLM cost per quality analysis

Photo by Max Petrunin on Unsplash

Price and score are the wrong two axes

Every launch this month came with the same two numbers: a per-token price and a benchmark. On their own, neither tells you what you want to know. A rate card of $1.40 per million looks cheap until the model burns three times the tokens to answer. A 60 on the intelligence index looks worth it until you see the model one point behind costs half as much to run. The question that actually maps to your invoice is the ratio: how much quality does each dollar buy?

There is a clean way to measure that, and it is not a rate card. Artificial Analysis publishes an Intelligence Index, a 0-to-100 composite across reasoning, coding, math, and knowledge evals, and alongside it a figure most people scroll past: the dollar cost to run that entire eval suite through each model's API. That number counts the real tokens spent, input and output, so it bakes in verbosity and reasoning overhead that a sticker price cannot show. It is a fixed basket of work priced through every model, which makes it the fairest apples-to-apples value metric we have.

So we pulled the current index score and the cost-to-run figure for eleven models people are actually choosing between in July 2026, straight off the Artificial Analysis pages on the 23rd. Then we did one thing the model pages do not: we sorted by value and asked which models a rational buyer would never pick.

The board, ranked by quality

Here is the full set, sorted by Intelligence Index from the top down. The last column is the one that reorders your intuition: cost to run the index divided by the score, or roughly what one point of measured quality costs on each model.

ModelAA IndexIn / Out /1MCost to run index$ / index point
Claude Fable 560$10 / $50$5,631$94
GPT-5.6 Sol59$5 / $30$2,824$48
Kimi K357$3 / $15$2,710$48
Claude Opus 4.856$5 / $25$3,753$67
GPT-5.6 Terra55$2.50 / $15$2,060$37
Grok 4.554$2 / $6$602$11
GLM-5.251$1.40 / $4.40$925$18
GPT-5.6 Luna51$1 / $6$870$17
Gemini 3.6 Flash50$1.50 / $7.50$727$15
DeepSeek V4-Pro44$0.435 / $0.87$176$4
DeepSeek V4-Flash40$0.14 / $0.28$74$2

Read the last two columns together and the shape of the market falls out. The top four scores sit inside five points of each other, 60 down to 56, but the cost to run them ranges from $2,710 to $5,631. You are not paying for quality up there. You are paying for the last sliver of it, and the price of that sliver is wildly uneven. One point of quality costs $94 on Fable 5, $67 on Opus 4.8, and $48 on both GPT-5.6 Sol and Kimi K3. Down at the bottom it costs $2. That spread is the whole story.

The value frontier, and the four models that fell off it

Borrow one idea from economics. A model sits on the efficient frontier if no other model beats it on both axes at once, higher score and lower cost. If some other model is both smarter and cheaper to run, the first model is dominated, and a buyer optimizing for value would never choose it. Plot our eleven and seven land on the frontier. Four do not.

ModelIndex / cost to runVerdict
Claude Fable 560 / $5,631On frontier (top score)
GPT-5.6 Sol59 / $2,824On frontier
Kimi K357 / $2,710On frontier
Claude Opus 4.856 / $3,753Dominated by Kimi K3 and Sol
GPT-5.6 Terra55 / $2,060On frontier
Grok 4.554 / $602On frontier
GLM-5.251 / $925Dominated by Grok 4.5 (API only)
GPT-5.6 Luna51 / $870Dominated by Grok 4.5
Gemini 3.6 Flash50 / $727Dominated by Grok 4.5
DeepSeek V4-Pro44 / $176On frontier
DeepSeek V4-Flash40 / $74On frontier (cheapest)

The headline is Opus 4.8. Anthropic's standard flagship, the model plenty of teams reach for by reflex, is now beaten outright by a Chinese model that launched a week ago. Kimi K3 scores 57 to Opus's 56 and runs the index for $2,710 against Opus's $3,753, so it wins on quality and undercuts on cost by 28% at the same time. GPT-5.6 Sol dominates Opus from the other side, three points higher for a thousand dollars less. When two separate models beat you on both axes, the rate card is telling you something.

The quieter story is Grok 4.5. It scores 54 for $602, and that single position knocks three models out at once. Gemini 3.6 Flash (50), GLM-5.2 (51), and GPT-5.6 Luna (51) all score lower and cost more to run than Grok does. Google shipped 3.6 Flash two days ago as a cost-focused release, and on this measure it is already dominated by a model that has been out since July 8. Cheap sticker price did not save it, because Grok answers the same eval for less actual money.

What Fable 5's last point actually costs

Fable 5 is not dominated. It holds the top score at 60, and if you genuinely need the single most capable model that money can rent, it is the one. But look at what the top of the ladder costs. Going from GPT-5.6 Sol at 59 to Fable 5 at 60 is one index point. The cost to run the suite goes from $2,824 to $5,631, almost exactly double. You are paying a 99% premium for a 1.7% bump in measured quality.

That is not an argument against Fable 5. Frontier work sometimes turns on that last point, and a single failed generation in production can cost more than the whole month of premium. It is an argument for knowing which side of the trade you are on. If your workload does not visibly break on Sol or Kimi K3, the Fable 5 premium is buying you a number on a benchmark, not an outcome you can feel. The Fable 5 pricing write-up covers the credit billing that complicates its rate card even further.

The same board as a monthly bill

The cost-to-run figure is a standardized basket, useful for ranking but not your invoice. So here is the plainer lens: a steady month of 15 million input tokens and 5 million output tokens, priced on each model's list rate. This is the number that lands on a card.

ModelAA Index15M in / 5M out month
Claude Fable 560$400.00
GPT-5.6 Sol59$225.00
Claude Opus 4.856$200.00
Kimi K357$120.00
GPT-5.6 Terra55$112.50
Grok 4.554$60.00
Gemini 3.6 Flash50$60.00
GPT-5.6 Luna51$45.00
GLM-5.251$43.00
DeepSeek V4-Pro44$10.88
DeepSeek V4-Flash40$3.50

On the sticker-price lens Kimi K3 looks even better, $120 a month for a 57 while Opus 4.8 asks $200 for a 56. But this table hides something the cost-to-run figure catches, and it is worth saying out loud so nobody gets burned. Kimi K3 is verbose. Artificial Analysis logged it spending about 130 million tokens across their suite against a 63 million average, so on a real task it emits far more output tokens than a fixed 5M-per-month assumption implies. That is exactly why the cost-to-run number, which counts actual tokens, still puts Kimi at $2,710 rather than dead last. The good news is that even after its own verbosity is priced in, Kimi still beats Opus. The bad news is your 5M output budget will not stretch as far on Kimi as this row suggests.

The bottom of the table is where the ratio gets absurd. DeepSeek V4-Pro scores 44, about three-quarters of Fable 5's quality, for $10.88 a month against Fable's $400. That is 73% of the index for under 3% of the cost to run it. For a large class of work, summarization, extraction, classification, routing, drafting, a 44 is not a compromise, it is plenty, and paying frontier rates for it is just leakage. Run your own token mix through the calculator before you assume you need the top of the board.

Where this measure stops being fair

A single ratio cannot carry a whole buying decision, and I would rather flag the holes than let the tables oversell. Four things the cost-per-quality frame does not capture, and each one can flip a call.

First, the index moved under Opus 4.8. At its May launch Opus 4.8 scored 61.4 and ranked first, on an older version of the Intelligence Index. Artificial Analysis revised the index since, and on the current one Opus reads 56. The model did not get worse, the ruler changed, so do not compare a launch-day number against a July number and conclude a model regressed. Everything in these tables is measured on the same current index, which is the only way the comparison holds.

Second, the composite hides task shape. Opus 4.8 and Fable 5 lead on coding and agentic reliability in ways a blended score flattens, and Kimi K3's measured hallucination rate is higher than the Claude models'. If your work is a narrow slice, code review, long-horizon agents, low-tolerance extraction, the model that wins your slice may not be the model that wins the average. Benchmark the shortlist on your actual prompts.

Third, GLM-5.2 and both DeepSeek models are open weights, and the cost-to-run figure prices only their hosted API. Self-host them at scale and the economics change completely, which can pull GLM-5.2 back onto the frontier it sits just off on API pricing. The GLM-5.2 pricing piece gets into that trade.

Fourth, none of this is latency or reliability. Grok 4.5 looks magnificent on value, but a 500K context ceiling, peak-hour pricing on the Chinese models, and rate limits on a preview tier are all real constraints that never show up in a dollar-per-point column. Value ranks the menu. It does not place your order.

Buy the cheapest model that clears your bar

The rule the frontier suggests is simple: pick the cheapest model that clears your quality bar, not the smartest model you can afford. Those are different instructions and they usually point at different models. If you are on Opus 4.8 out of habit, the honest move this month is to run Kimi K3 and GPT-5.6 Sol against your evals, because both now beat it on paper and one of them costs less. If you are on Gemini 3.6 Flash or GPT-5.6 Luna for cost reasons, put Grok 4.5 next to them, because it scores higher for less.

And if you have never checked whether your cheap-tier traffic needs a 50 or would be fine on a 44, that one test is the hour of work that pays for itself fastest. The gap between DeepSeek V4-Pro and the mid-tier is a 10x cost difference for a dozen-point quality difference that most routine tasks will never notice.

The pricing page has every current rate card side by side, and the calculator runs your real token mix against all of them at once. For the model that started this reshuffle, see the Kimi K3 write-up, and for the flagship it passed, the Opus 4.8 breakdown.

Sources

A note on benchmarks: SWE-bench Verified and Pro figures circulating for several of these models are vendor-reported or from aggregators, and OpenAI and xAI did not publish SWE-bench Verified for GPT-5.6 or Grok 4.5 at all. We deliberately built this analysis on the Artificial Analysis Intelligence Index and its cost-to-run figure, which are measured on one consistent methodology, rather than on contested per-vendor coding scores.