Skip to main content
TokenCost logoTokenCost
ResearchAugust 17, 2026·11 min read

Terminal-Bench 3.0 publishes what every run cost, which almost no other leaderboard does. Two of its ten rows are billed at a card OpenAI retired on July 30, and repricing them makes the ninth-place model the cheapest per solved task by 3.5x.

Benchmark leaderboards almost never tell you what the run cost. This one has a COST column and a TOKENS column sitting right next to the score, broken down per row into output, prompt and cached tokens, to the cent. That is a genuinely useful thing to publish and we would like more of it. It also creates a trap, because a cost column is a record of what something cost on the day it ran, and readers treat it as a price list. Ten rows, run over several weeks, priced at whatever each vendor was charging that week. We reproduced four of the ten figures exactly from the published token counts, and the four that reproduce tell you precisely which card each run was billed at. Two of them are billed at rates that no longer exist.

A worn brass mechanical tally counter on industrial machinery, reading over five million

Photo by Alexandre Daoust on Unsplash

What one solved task costs

Run cost divided by tasks solved, where tasks solved is 370 trials times the resolution rate. The two lighter bars are our repricing of the OpenAI rows at the card in force today. Read on for why only those two move.

GPT-5.6 Luna, repriced
$6.04
Grok 4.6
$21.37
GPT-5.6 Terra, repriced
$25.83
GPT-5.6 Luna, as published
$30.19
GPT-5.6 Sol
$30.86
GPT-5.6 Terra, as published
$32.28
Opus 5
$36.83
Fable 5
$51.37
Opus 4.8
$66.79
Sonnet 5
$127.56
GLM-5.2
$199.87

Grok 4.5 is omitted from this chart for the same reason the benchmark authors omit it from theirs: they flag a cost reporting issue with the Cursor CLI harness on that row. Its $766.02 would otherwise rank first at $13.19.

The column almost nobody else publishes

Terminal-Bench 3.0 went up on Harbor Hub on August 7 as revision 1 of a 74 task dataset, and the board there carries columns for rank, agent, model, reasoning effort, accuracy, release date, agent org, model org, tokens and cost. Scores are a mean reward with confidence intervals between 1.0 and 1.7 percentage points, and the agent and the verifier run in separate containers so that artifacts can be re-graded later. We loaded the board on August 17 and every figure below is from that snapshot.

One number we use throughout deserves flagging up front, because every cost-per-solved-task figure in this post divides by it. We treat each run as 370 trials, that is 74 tasks attempted five times each. The 74 is the benchmark's own, confirmed on Harbor Hub. The five attempts per task are not stated anywhere on the leaderboard or the dataset page; the figure comes from a footnote in Anthropic's Opus 5 announcement describing its own run as a mean reward over five attempts per task. If some rows used a different number of attempts, the absolute dollar figures per solved task would move while the ordering would not, since the denominator scales every row alike.

One correction worth making early, because it is easy to repeat second-hand. Terminal-Bench 3.0 is not Scale's benchmark. It comes from the Harbor and Laude Institute team, it was called Frontier-Bench until recently, and Scale is one of about ten sponsors alongside Modal, Anthropic, OpenAI and Google. Scale blogged about it, which is how the attribution drifted. The rename also explains a small mystery: Anthropic's Opus 5 materials cite "Frontier-Bench v0.1" and never use the words Terminal-Bench, so anyone grepping the Opus 5 announcement for a Terminal-Bench score comes away empty even though the score is right there. Anthropic reports 43.3% from its own internal run against the official board's 42.7%.

#ModelAgentScoreTokensCostPer solved
1Opus 5mini-SWE-agent42.7%7.28B$5,818.20$36.83
2GPT-5.6 SolCodex34.6%5.76B$3,950.90$30.86
3Fable 5Claude Code34.1%3.58B$6,480.94$51.37
4Grok 4.6Grok Build26.5%2.85B$2,095.44$21.37
5Opus 4.8Claude Code21.1%5.21B$5,214.05$66.79
6GPT-5.6 TerraCodex20.8%7.02B$2,484.52$32.28
7Grok 4.5Cursor CLI15.7%1.24B$766.02$13.19
8Sonnet 5Claude Code14.6%17.95B$6,890.54$127.56
9GPT-5.6 LunaCodex14.3%11.87B$1,597.50$30.19
10GLM-5.2Claude Code4.6%3.28B$3,401.76$199.87

The per solved column is ours. So is the precision: the public board rounds to $5.8k and 7.3B, and only Grok 4.5's $766.02 appears there to the cent, so the exact costs and token counts above come from the underlying run data rather than the rendered table. Note what the last column does to the ordering: Opus 5 wins the benchmark by 8.1 points and comes sixth on cost efficiency, while Grok 4.6 sits fourth on score and first among rows we trust on cost. Note also that GLM-5.2 is not merely last, it is an order of magnitude off the pace, because a 4.6% resolution rate divides a $3,401.76 bill across only 17 solved trials.

Four of the ten cost figures reproduce to the cent

Each row page publishes the run's token usage split three ways into output, prompt and cached. That is enough to test the cost column against a rate card. Multiply each leg by the matching published price, add them up, and see whether you land on the number in the COST column. For four rows you land on it exactly, which is not something that happens by accident at this many significant figures.

RowCard that reproduces itComputedPublishedStill current?
Opus 5$5.00 in, $0.50 cached, $25.00 out$5,818.20$5,818.20Yes
GPT-5.6 Sol$5.00 in, $0.50 cached, $30.00 out$3,950.90$3,950.90Yes
GPT-5.6 Terra$2.50 in, $0.25 cached, $15.00 out$2,484.52$2,484.52No, cut July 30
GPT-5.6 Luna$1.00 in, $0.10 cached, $6.00 out$1,597.50$1,597.50No, cut July 30

The pattern inside those four rows is the whole finding. All three Codex runs carry a July 9 release date, so all three ran before OpenAI's July 30 repricing. Sol's card did not move that day and its row reproduces at the price you would pay this morning. Terra and Luna were cut, and their rows reproduce only at the pre-cut card: Terra at $2.50 and $15.00 rather than $2.00 and $12.00, Luna at $1.00 and $6.00 rather than $0.20 and $1.20. We covered that cut when it happened in our July 31 post, which is where the pre-cut figures come from.

We want to be careful about what this is and is not. It is not an error. A benchmark that records what a run cost on the day it ran is doing the honest thing, and repricing history every time a vendor moves would make the record less reproducible rather than more. The problem is entirely on the reading side. A column headed COST, sat beside a live score, invites you to compare vendors on it today, and for two of these ten rows that comparison is against prices that were withdrawn eighteen days ago.

The remaining six rows do not reproduce under the same formula and we are not going to pretend otherwise. The two Grok rows sit between the short-context and long-context bounds of xAI's card, which is what you would expect when some requests cross the 200K threshold and bill at double. The four Claude Code rows miss in both directions, which rules out the tidy explanation. Sonnet 5 and GLM-5.2 bill above the floor the three published legs imply, by $1,730.27 and $2,129.49, and cache writes would account for that. But Fable 5 and Opus 4.8 bill $1,562.05 and $263.02 below their floors, and nothing billed on top of output, prompt and cached tokens can make a total smaller. Something in those rows is either discounted or counted differently from what the split shows, and we do not know which. For all six we can tell you what the run cost and not which card produced it.

Repricing moves ninth place to first

OpenAI cut every line on Luna's card by exactly 80% and every line on Terra's by exactly 20%. Because the cut is uniform across input, cached and output, the repricing is a single multiplication rather than a re-derivation: Luna's run costs exactly one fifth of what the board says, and Terra's exactly four fifths.

RowBoard saysAt today's cardPer solved, beforePer solved, after
GPT-5.6 Luna$1,597.50$319.50$30.19$6.04
GPT-5.6 Terra$2,484.52$1,987.62$32.28$25.83

At $6.04 per solved task Luna finishes the same work for 28 cents on the dollar against Grok 4.6, the cheapest row we trust as published, and it does that from ninth place on the board. The obvious objection is that a model solving 14.3% of tasks is not interchangeable with one solving 26.5%, and that objection is correct. Cost per solved task is a ratio, and a low ratio earned at a low resolution rate tells you the model is efficient at the subset it can do, not that it can do the rest. Luna is the right lens if your work resembles the easy 14% and the wrong lens if it resembles the hard 60%.

One thing we checked before publishing this, because it would invalidate the whole comparison if it were false: whether any other row's card has moved since its run. Anthropic's Opus 5, Opus 4.8, Fable 5 and Sonnet 5 rates are unchanged, and Sonnet 5's $2 and $10 is now permanent after Anthropic cancelled the September 1 increase we wrote about in July. Grok 4.6 kept Grok 4.5's card except for the cached-input line, and Grok 4.6's run postdates its own launch. Z.ai's GLM-5.2 card is unchanged. Of ten rows, exactly two are priced at rates you can no longer buy, and both belong to OpenAI.

This is a cache-read benchmark wearing a terminal costume

Token splits on this board are lopsided in a way that should change which number you look at first on a rate card. Cached input is between 92.4% and 98.2% of all tokens on every single one of the ten runs. Generated output, the line most pricing coverage leads with because it carries the biggest sticker, is between 0.37% and 1.85%. An agent loop replays its context on every turn, so what you are mostly buying is the same tokens back again.

RowCached share of tokensOutput share of tokensCache reads as share of bill
GPT-5.6 Luna98.24%0.37%73.0%
GPT-5.6 Sol98.08%0.40%71.5%
GPT-5.6 Terra97.85%0.44%69.2%
Opus 597.40%0.91%60.9%

That last column covers only the rows whose card we could confirm by reproduction, because computing a share of the bill requires knowing the prices that produced it, and those are the four rows we could reproduce. Within them the cached-input line is the single largest component every time, running from 60.9% on Opus 5 to 73.0% on Luna. The token shares in the first two columns need no such caveat, since they are counts rather than money, and there the pattern holds across all ten rows without exception. That is the practical takeaway from this whole exercise and it costs nothing to act on: when you evaluate a model for agent work, price the cached-input line first and the output line last. It is also why Grok 4.6 raising only its cache rate, which we wrote about on August 15, was a bigger change than it looked.

Sonnet 5 spent 17.9 billion tokens to finish 54 trials

The most expensive run on the board is not the most expensive model. Sonnet 5 lists at $2 and $10, a fifth of Fable 5 and 40% of Opus 5, and it produced the largest bill of the ten at $6,890.54. It got there by consuming 17.95 billion tokens, which is 2.47 times what Opus 5 used, for a score roughly a third of Opus 5's. Cheaper per token, dramatically more expensive per outcome, at $127.56 against $36.83.

There is a caveat we should attach rather than bury. Anthropic moved to a new tokenizer with Claude 4.7 that produces roughly 30% more tokens for the same text, which we measured against production bills in May. Raw token counts are therefore not comparable across vendors, and some of the gap between 17.95B and 11.87B is a counting convention rather than a workload difference. Dollars are comparable, though, because each vendor bills its own tokens at its own rates, and in dollars Sonnet 5 finished second from last.

Do not put a 2.1 score in the same table as a 3.0 score

Almost every vendor launch page still quotes Terminal-Bench 2.x, and those numbers look far better because 2.1 is saturated. The authors built 3.0 as a harder replacement and say so in their announcement, which is also where they set out the container separation and the re-grading design. Grok 4.5 makes the gap impossible to miss: xAI reports 83.3% on Terminal Bench 2.1, and the official 3.0 board shows the same model at 15.7%. That is 67.6 points apart on one model from one vendor, and the 83.3% is xAI's own published figure. OpenAI's GPT-5.6 launch materials are reported as putting Sol at 88.8 on 2.1 against 34.6 here; the corresponding 2.1 figures for Terra and Luna are cited by trackers as 87.4 and 84.7, but OpenAI's page renders its benchmark charts as images we could not read directly, and sources disagree on Luna, so treat those two as secondhand.

Harness choice moves scores too, and by more than most tables admit. Anthropic's own Opus 4.8 materials concede a 5.2 point swing on an identical model between two harnesses, in that case on GPT-5.5 rather than on its own model. On this board every row names its agent, and five different harnesses appear across ten rows, so some of what reads as a model gap is a harness gap. Opus 5 tops the board running mini-SWE-agent, a Princeton harness, while Anthropic's own Claude Code sits behind it on all four rows where it appears, which is a result worth sitting with for a moment.

Read it as a receipt, not a quote

Treat the cost column as a dated receipt, not a quote. Check each row's run date against the vendor's pricing history before you put two rows side by side, because right now two of these ten are eighteen days stale and the board has no way to tell you that. If you want the shortcut: multiply the Luna row by 0.2 and the Terra row by 0.8, and leave the other eight alone.

Then stop reading cost per token and start reading cost per finished job. The spread on this board is 9.35x from Grok 4.6 to GLM-5.2 on cost per solved task, and 33.1x if you use our repriced Luna as the floor. Sticker price will not get you to that ordering: GLM-5.2 lists at $1.40 input against Sol's $5.00, and still costs 6.5 times more per solved task, $199.87 against $30.86. The ranking by sticker price and the ranking by cost to get work done are close to unrelated, and only one of them shows up on your invoice. Our cost calculator takes a cached, input and output split like the ones above, and current rates for every model named here are on the pricing page.

What each page here can and cannot prove

  • Terminal-Bench 3.0 leaderboard - Read on August 17, 2026. All ten rows: rank, model, agent, resolution rate with confidence interval, release date, rounded token total and rounded cost, plus the note that the site was formerly called Frontier-Bench. The Grok 4.5 caveat is not on this page: it is footnote one on the separate announcement page, attached to the Cost vs Pass Rate chart, and reads "Grok 4.5 was removed from this chart due to a cost reporting issue with the Cursor CLI harness". The announcement is also where the authors describe separating the agent and verifier containers to allow re-grading, and where they say many Terminal-Bench tasks have become saturated
  • Harbor Hub: terminal-bench dataset - The canonical source the leaderboard mirrors, published August 7, 2026 by Alex Shaw as revision 1 at version 3.0.0, showing 74 of 74 tasks and the same ten rows, with agent org and model org columns the public board omits. The exact costs and the three-way output, prompt and cached token splits that every calculation here is built on come from the underlying run data rather than a page we can link you to, so treat them as the weakest link in the chain. Two independent checks support them: all ten splits sum to the rounded token totals on the public board, and four of them reproduce the published costs to the cent at published rate cards, which is not something wrong numbers do twice
  • OpenAI: API pricing - Read August 17, 2026. GPT-5.6 Sol at $5.00 input, $0.50 cached and $30.00 output; Terra at $2.00, $0.20 and $12.00; Luna at $0.20, $0.02 and $1.20. These are the rates behind the repriced column, and Sol's are the rates that reproduce its row exactly
  • TokenCost: OpenAI cut GPT-5.6 Luna by 80% and left Sol untouched - Our July 31 coverage of the July 30 repricing, and the source of the pre-cut card: Luna at $1.00 and $6.00 with every line down exactly 80%, Terra from $2.50 and $15.00 down 20% to $2.00 and $12.00, Sol unchanged. Those pre-cut figures are what reproduce the Terra and Luna rows to the cent
  • Anthropic: model pricing - Opus 5 and Opus 4.8 at $5.00 input, $0.50 cache read and $25.00 output; Fable 5 at $10.00, $1.00 and $50.00; Sonnet 5 at $2.00, $0.20 and $10.00, now permanent after the September 1 increase was cancelled. Opus 5's figures reproduce its row exactly; the Claude Code rows do not reproduce under the same formula
  • xAI: pricing - Grok 4.6 at $2.00 input, $0.50 cached and $6.00 output below 200K context and double above it, Grok 4.5 identical except $0.30 cached. Note that the separate docs.x.ai models page omits cache rates entirely, so this is the page to cite. The doubling above 200K is why the two Grok rows fall between our short-context and long-context bounds
  • Z.ai: pricing - GLM-5.2 at $1.40 input, $0.26 cached and $4.40 output, unchanged since its June launch, which is what lets us say its row is not stale even though we cannot reproduce it
  • TokenCost: Grok 4.6 kept Grok 4.5's card except the cache line - Our August 15 post on the cached-input rise from $0.30 to $0.50, and on why a cache-only price change hits agent workloads hardest. The 92% to 98% cached-token shares measured here are the direct evidence for that argument
  • TokenCost: the Opus 4.7 tokenizer change in production - Where the roughly 30% more tokens per unit of text figure comes from, and the reason we compare these runs in dollars rather than in raw token counts
  • Four limits on this analysis, stated plainly. Six of the ten cost figures do not reproduce under the output-plus-prompt-plus-cached formula, so for the four Claude Code rows and the two Grok rows we report the published cost without being able to say which card produced it, and the cache-read share column is restricted to the four rows we could verify. Cost per solved task is our own arithmetic, not the benchmark's, and it flatters models that score low, because dividing by a small number of solved trials rewards efficiency on the easy subset rather than breadth. The 370-trial denominator under every one of those figures rests on five attempts per task, which the benchmark does not publish and which we took from Anthropic's description of its own run. And the leaderboard's date column is the model's release date rather than the eval date, so our claim that the three Codex runs predate July 30 rests on that release date plus the fact that their costs reproduce at the pre-July-30 card and not at today's, rather than on any published eval timestamp