Grok 4.5 sells the cheapest token, Fable 5 sells the smartest one, and neither is the cheapest way to finish the actual job. You only see that once you stop pricing tokens and start pricing benchmark points.
Three frontier models landed inside a month: GPT-5.6 Sol on July 9, Grok 4.5 the day before, and Claude Fable 5 back in circulation from June. Their rate cards span five times over, from $2/$6 to $10/$50. On a per-token basis the ranking writes itself. Divide each price by what the model actually scores and the order scrambles, because a cheap token you need more of, or one that scores lower when you get there, was never really cheap. Here is the cost-per-capability math on all three.

Photo by Aakash Dhage on Unsplash
Three sticker prices, a fivefold spread
| Model | Input / output per 1M | Context |
|---|---|---|
| Grok 4.5 | $2 / $6 | 500K |
| GPT-5.6 Sol | $5 / $30 | ~1M |
| Claude Fable 5 | $10 / $50 | 1M |
USD per million tokens. Grok caches input at $0.50 and bills above 200K tokens at a higher rate that secondary trackers put near double; Fable reads cache at roughly $1 and runs Batch at half. Sources: OpenAI, Anthropic, and xAI docs plus Artificial Analysis, which had pre-release access to all three.
Per token, Grok wins and it is not close
Start with the easy read. Grok 4.5 costs 60 percent less than Sol on input and five times less on output. Against Fable it is five times cheaper going in and more than eight times cheaper coming out. If the only number on your dashboard were dollars per million tokens, you would close this tab and move your traffic to Grok this afternoon. Plenty of teams will, and for a lot of workloads that is the correct call.
But nobody buys tokens to own tokens. You buy them to get a pull request merged, a ticket closed, a document summarized correctly the first time. The token is an input to the work, and the price that matters is the price of the work. Two models charging the same per token can cost wildly different amounts per finished task if one of them needs twice as many tokens or fails half as often. Which means the per-token table is where the analysis starts, not where it ends.
The capability gap is smaller than the price gap
Here is what the fivefold price spread is actually buying you in capability. Artificial Analysis, which evaluated all three on its v4.1 Intelligence Index, has them bunched near the top: Fable 5 at 60, Sol at 59, Grok at 54. On raw coding the picture splits. Sol and Grok land in a dead heat in the mid-60s on SWE-bench Pro, at least as the independent trackers report it, while Fable pulls clear at roughly 80 percent, the widest single gap in the set. Sol takes the coding-agent harness back by a nose.
| Benchmark | Grok 4.5 | GPT-5.6 Sol | Fable 5 |
|---|---|---|---|
| AA Intelligence Index v4.1 | 54 | 59 | 60 |
| SWE-bench Pro | 64.7% | 64.6% | ~80% |
| AA Coding Agent Index | 76 | 80 | 77.2 |
Sit with that a second. The price gap between Grok and Fable is 500 percent. The intelligence gap is six points, or about 11 percent. Even the big SWE-bench Pro gap, the one place Fable genuinely separates, is a quarter, not a multiple. Capability is compressing at the top while price still fans out five to one. That is the tension the rest of this post is about, and the cleanest way to cut it is to put the two numbers in the same fraction.
Cost per benchmark point
Pick one representative task you can price by hand: a large agentic coding run, 500K tokens of input (a chunk of a repo plus history) and 50K of output (diffs, plans, explanations). Multiply by each rate card, then divide by the SWE-bench Pro score. The result is dollars per point, a rough proxy for what a unit of coding capability costs you on each model.
| Model | Task cost (sticker) | SWE-bench Pro | $ per point |
|---|---|---|---|
| Grok 4.5 | $1.30 | 64.7 | $0.020 |
| GPT-5.6 Sol | $4.00 | 64.6 | $0.062 |
| Claude Fable 5 | $7.50 | 80.0 | $0.094 |
On this measure Grok is about three times cheaper per point than Sol and nearly five times cheaper than Fable. Sol and Grok score within a tenth of a point of each other on SWE-bench Pro, so the entire difference between them here is price. Fable earns its 80, but you pay for every one of those extra 15 points at Fable rates, and the per-point number punishes that hard.
One honest caveat, because this metric can mislead if you let it. A benchmark percentage is not a linear currency. Going from 64 to 80 on SWE-bench Pro is not "25 percent more work done," it is a harder class of ticket getting solved at all, and for some teams that top slice is the only slice that pays. Cost per point rewards the cheap floor and taxes the expensive ceiling, which is exactly right if your work lives in the middle and exactly wrong if you live at the edge. Know which you are before you let the fraction decide.
The twist: these models do not spend tokens equally
The table above hands every model the same 50K output budget. In reality they do not use it the same way, and that is where sticker math and measured math part company. Artificial Analysis ran its own harness with real token counts and priced a task on each. Sol came in near $1.04, Grok near $0.31, and it describes Fable as running roughly three times Sol, so about $3. The ordering holds, Grok cheapest and Fable priciest, but look at the middle.
| Model | Measured cost / task | What you get for it |
|---|---|---|
| Grok 4.5 | ~$0.31 | Index 54, mid-60s coding, the outright floor |
| GPT-5.6 Sol | ~$1.04 | Index 59, one point under the leader |
| Claude Fable 5 | ~$3 | Index 60 and ~80 on SWE-bench Pro |
Sol is the interesting line. It scores 59 on the Intelligence Index, a single point below Fable's 60, and Artificial Analysis measures it doing so for about a third of Fable's per-task cost. That is the most defensible value claim in the whole trio: if what you want is a top-two intelligence score, Sol is the cheap way to buy one. Grok gives up five index points to save money below even Sol; Fable spends the most of the three to win the index by a single point. Three models, three different bargains, and the token rate told you almost none of it.
Pick the corner, not the winner
There is no single answer here, and any post that hands you one is selling something. What there is instead is three clean corners, and your workload already sits in one of them. If the invoice is the binding constraint and mid-60s SWE-bench Pro clears your bar, Grok 4.5 is the pick and it is not particularly close: cheapest per token, cheapest per point, cheapest per measured task. You give up five points of index and half your context window, and for a great deal of agent and coding work you will not miss either.
If your work lives at the top of SWE-bench Pro, the 80 percent tier where the hard tickets get closed, Fable 5 is the only one of the three that reaches it, and you pay Fable rates for the privilege. That is a real product, not a rip-off, as long as you are honestly at the edge and not just buying the biggest number to feel safe. And if you want the best intelligence score per dollar without dropping to the budget floor, Sol is the answer the measured numbers keep pointing at, one point off the lead for a third of the price.
All of this rides on sticker rates and published benchmarks, and your real bill turns on your own token mix and how often each model needs a second try on your tasks. Drop your actual input-output split into the cost calculator and line the three up on the pricing page. The per-point framing tells you where to look; only your own workload tells you where to land.
Sources
- - Artificial Analysis, GPT-5.6 evaluation: artificialanalysis.ai
- - Artificial Analysis, Grok 4.5 evaluation: artificialanalysis.ai
- - Anthropic, Claude Fable 5 docs: platform.claude.com
- - Simon Willison on GPT-5.6: simonwillison.net
- - Related: Grok 4.5 pricing breakdown
- - Related: Claude Fable 5 metered billing
- - TokenCost pricing page: tokencost.app/pricing