Every model measured across both rewrites of the industry's most-quoted leaderboard lost between 12.3 and 18.1 points in four days, and not one of them changed. The ruler did, twice, and it moved cost per point by 154% for one model and 290% for another.
Artificial Analysis published Intelligence Index v4.2 on September 4 and v4.3 on September 7. Claude Fable 5.1 scored 65.65 before them and 53.37 after. No lab shipped weights in between, and not one rate card on this page moved a cent. What moved was the measuring instrument, and because it also got longer to run, the cost of being measured by it roughly doubled at the same time.

Photo by Spencer Liao on Unsplash
What moved, and what did not
Nothing about any model changed this week. No retraining, no new checkpoint, no repricing. Every number below moved because Artificial Analysis rewrote its Intelligence Index on September 4 and again on September 7, and the version in between was current for three days.
What that did to scores is uniform in direction and nothing like uniform in size: every model on the board before September 4 fell between 12.3 and 18.1 points, by a different amount each. What it did to cost is the half that belongs on a pricing site. The last version step alone made running the index 25% to 103% dearer per model on rate cards that did not move, because the suite got longer and models now emit half again to twice the output tokens per task. Claude Fable 5.1 went from $3.69 a task to $7.63 in five days while listing at $10.00 and $50.00 per million throughout.
Divide the two and cost per intelligence point rose by between 154% and 290% across the models measured through both rewrites, which is not a recalibration anyone can divide back out. Rank eight models by that figure and the middle four change places with nothing but the methodology moving. Artificial Analysis writes "not directly comparable" four times on its methodology page about individual benchmarks, and has never written it about the composite those benchmarks feed. Our own rankings are three revisions stale and carry no version stamp at all, which we get to below.
The version that was current for three days
Artificial Analysis keeps a version history on its methodology page with an active window for each revision. Read it in order and the cadence is the story before any score is.
| Version | Published | Window AA states | Time in force |
|---|---|---|---|
| v4.1 | June 15, 2026 | June 2026 to August 2026 | about 7 weeks |
| v4.1.1 | August 6, 2026 | August 2026 to September 2026 | 4 weeks |
| v4.2 | September 4, 2026 | September 2026 to September 2026 | 3 days |
| v4.3 | September 7, 2026 | September 2026 to current | live now |
We pinned the switchover with the Wayback Machine rather than trusting the article dates, because the leaderboard prints its own version in a line reading "Intelligence Index v4.X incorporates N evaluations". The September 3 capture says v4.1.1. The September 5 capture says v4.2, and so does the capture at 03:03 UTC on September 7. The live page today says v4.3. So v4.3 went live at some point after 03:03 on the seventh, and anybody who pulled the leaderboard that morning got a number that had hours left on it.
There is a third change already announced. The v4.3 article's standfirst closes on the sentence "This is a continuation of our rollout of Intelligence Index v5." Whatever number you write down this week has a known replacement coming.
A quarter of the index by weight was replaced, then a tenth of it was replaced again
The coverage this week has focused on v4.3, because that is the one that produced the tie at the top. By weight, v4.2 was the much larger rewrite: it introduced two benchmarks nobody had been scored on before, worth 25% of the total between them, and deleted GPQA Diamond outright.
| Evaluation | v4.1 | v4.2 | v4.3 |
|---|---|---|---|
| GDPval-AA v2 | 20% | 10% | 10% |
| Terminal-Bench | 16% (v2.1) | 10% (v2.1) | 10% (v4.0) |
| τ³-Banking | 14% | 5% | removed |
| AutomationBench-AA | not present | not present | 5% |
| AA-Briefcase | not present | 15% | 15% |
| GDP.pdf | not present | 10% | 10% |
| HLE | 12% | 10% | 10% |
| CritPt | 6% | 10% | 10% |
| SciCode | 8% | 10% | 10% |
| AA-LCR | 6% | 5% | 5% |
| Omniscience Accuracy | 8% | 10% | 10% |
| Omniscience Non-Hallucination | 4% | 5% | 5% |
| GPQA Diamond | 6% | removed | removed |
The share of the index held in private test sets went 20% under v4.1, to 40% under v4.2, to 45% under v4.3. Artificial Analysis puts the middle step plainly: 40% of the weighting is now private, held-out test sets, which it notes is double the figure from v4.1. The v4.3 increment comes entirely from running AutomationBench-AA against a held-out set of 657 tasks in collaboration with Zapier. That is a defensible direction of travel. It also means a growing share of the number cannot be independently reproduced by anyone outside Artificial Analysis, which is worth holding in mind when the number moves.
One removal got a defect named and one did not. GPQA Diamond was cut as "an exceptional scientific reasoning evaluation that has now been saturated". τ³-Banking got no equivalent. What it got was a replacement rationale, that AutomationBench-AA covers broader business workflows and adds a private test set, which is an argument for the new benchmark rather than anything about the old one. Worth noting because Artificial Analysis had upgraded that benchmark on August 6, then cut its weight from 14% to 5% on September 4, then deleted it on September 7. Three changes to one slot in 32 days, one of them explained.
Everything fell, and the amount it fell by is different for every model
These are the exact values Artificial Analysis embeds in the structured data on its own leaderboard, read from the live page for v4.3 and from Wayback captures of the same URL for the two earlier versions. They are not chart estimates. The first column is the September 3 capture, which labels itself v4.1.1: a point release that changed grader models and the τ³ dataset while keeping v4.1's weights, so it is the last state of the board before the two September rewrites rather than v4.1 as first published.
| Model | v4.1.1 | v4.2 | v4.3 | Change |
|---|---|---|---|---|
| Claude Fable 5.1 (max) | 65.65 | 56.76 | 53.37 | -12.3 |
| Claude Opus 5 (max) | 63.05 | 54.05 | 50.70 | -12.4 |
| Muse Spark 1.3 (max) | 62.09 | 52.95 | 48.17 | -13.9 |
| Claude Fable 5 | 62.07 | 53.19 | 49.70 | -12.4 |
| GPT-6 Astra (max) | did not exist | 54.66 | 52.81 | -1.8 |
| GPT-5.6 Sol (max) | 60.93 | 51.26 | 47.06 | -13.9 |
| Grok 4.6 (high) | 60.92 | 50.58 | 44.41 | -16.5 |
| Kimi K3 (max) | 59.70 | 50.23 | 43.78 | -15.9 |
| GLM-5.3 (max) | 59.51 | 48.58 | 44.86 | -14.7 |
| Gemini 3.8 Flash (high) | 58.68 | 47.07 | 41.19 | -17.5 |
| GLM-5.3-Flash | 57.46 | 46.22 | 41.91 | -15.6 |
| GPT-5.6 Terra (max) | 56.58 | 46.77 | 42.25 | -14.3 |
| DeepSeek V4 Pro 0813 (max) | 53.20 | 42.11 | 36.28 | -16.9 |
| Qwen3.8 27B (xhigh) | 52.03 | 41.41 | 33.90 | -18.1 |
| MiniMax-M3 | 45.40 | 35.75 | 29.61 | -15.8 |
| Gemini 3.5 Flash-Lite | 37.44 | 27.57 | 22.66 | -14.8 |
The uneven part is the part that matters. Take just the v4.2 to v4.3 step and express it as a multiplier on the old score: Qwen3.8 27B comes out at 0.82, Kimi K3 at 0.87, Claude Fable 5.1 at 0.94, and GPT-6 Astra at 0.97. Across the sixteen models above, one version step multiplies scores by anything from 0.82 to 0.97, and the floor drops to 0.74 once you include the rest of the board. The weaker the model, the harder it was hit, so the board did not shift down so much as stretch. Fable 5.1 led Gemini 3.5 Flash-Lite by 28.21 points under v4.1.1 and leads it by 30.71 under v4.3, a wider gap between two models that both got smaller numbers.
GPT-6 Astra is worth a line of its own, and not as an accusation. Astra lost the least of any model on the board in the version that produced its tie with Fable 5.1: it kept 96.6% of its v4.2 score while Fable 5.1 kept 94.0%. Under v4.2 the two were 2.101 points apart. Under v4.3 they are 0.560 apart, which is how you get the headline that they both score 53. That gap closed without either model being touched. Astra also has three published index scores inside four days: Artificial Analysis' own launch piece on September 3 says it "scores equal to GPT-5.6 Sol in the Index at 61". Today it is 52.81 and Sol is 47.06, so the model that was reported as level with Sol is now 5.75 points clear of it. That September 3 article still says 61 and carries no correction.
The rate cards did not move. The bills did.
This is the part that belongs on a pricing site. Artificial Analysis publishes what it cost to run each model through the index, which almost no leaderboard does. Between September 3 and September 8 that figure rose for every model on the board, and not one of the prices in the second column changed in that window. The right-hand column isolates the last version step on its own; over both steps the same models run from +98% to +266%.
| Model | Rate card, in / out per 1M | $/task v4.1.1 | $/task v4.2 | $/task v4.3 | v4.2 to v4.3 |
|---|---|---|---|---|---|
| Claude Fable 5.1 (max) | $10.00 / $50.00 | $3.69 | $6.12 | $7.63 | +25% |
| Claude Opus 5 (max) | $5.00 / $25.00 | $2.34 | $4.21 | $5.86 | +39% |
| GPT-6 Astra (max) | $10.00 / $50.00 | did not exist | $2.57 | $3.26 | +27% |
| Kimi K3 (max) | not published here | $0.84 | $1.58 | $2.00 | +27% |
| GLM-5.3 (max) | $1.40 / $4.40 | $0.68 | $1.26 | $2.01 | +59% |
| GPT-5.6 Sol (max) | $4.00 / $20.00 | $0.95 | $1.25 | $1.99 | +59% |
| Grok 4.6 (high) | not published here | $0.94 | $1.25 | $1.86 | +48% |
| GPT-5.6 Terra (max) | $2.00 / $12.00 | $0.53 | $0.81 | $1.40 | +72% |
| Gemini 3.8 Flash (high) | $0.75 / $3.75 | $0.58 | $0.74 | $1.24 | +68% |
| DeepSeek V4 Pro (max) | not published here | $0.27 | $0.33 | $0.67 | +103% |
| GLM-5.3-Flash | $0.15 / $0.50 list | $0.09 | $0.18 | $0.25 | +38% |
| GPT-5.6 Luna (max) | $0.20 / $1.20 | $0.05 | $0.10 | $0.18 | +84% |
Fable 5.1 doubled in five days. $3.69 per task on September 3, $7.63 on September 8, on an unbroken $10.00 and $50.00 card. Running the full index against it now costs $13,128.86 against Astra's $5,324.10.
The mechanism is output tokens, not rates. Claude Fable 5 went from 35,565 output tokens per index task to 66,848, GPT-5.6 Terra from 20,838 to 38,897, Kimi K3 from 25,474 to 48,455. That is a shade under a doubling in each case. Gemini 3.8 Flash is the mildest of the four at 48,315 to 71,003, up about half. The tasks got longer and more agentic, and the token count followed. If you have been reading "cost per task" as a property of the model, it is at least as much a property of the suite, and the suite is now a different suite.
One caution on the cheapest row. Artificial Analysis prices GLM-5.3-Flash at $0.25 per task using Z.AI's list card of $0.15 and $0.50. Z.AI is currently selling it at half that under a launch promotion, and that promotion expires tomorrow, September 9 at 16:00 UTC. So the $0.25 is the post-promotion number and it is the right one to quote from Wednesday onward. If you compute it yourself from today's Z.AI pricing page you will get roughly half, and it will be wrong by Wednesday afternoon.
Divide one moving number by another moving number
Cost per intelligence point is the figure that gets quoted in procurement decks, and it is the ratio of the two things that both just moved: cost per task went up, index score went down. Neither Artificial Analysis nor anyone else publishes this series, so we computed it as $/task divided by index score for each version.
| Model | $ per point, v4.1.1 | $ per point, v4.3 | Change |
|---|---|---|---|
| Claude Fable 5.1 (max) | $0.0562 | $0.1430 | +154% |
| Claude Opus 5 (max) | $0.0371 | $0.1155 | +212% |
| GPT-6 Astra (max) | did not exist | $0.0617 | +31%, one step only |
| GPT-5.6 Sol (max) | $0.0156 | $0.0422 | +170% |
| Grok 4.6 (high) | $0.0154 | $0.0419 | +172% |
| Kimi K3 (max) | $0.0140 | $0.0457 | +226% |
| GLM-5.3 (max) | $0.0115 | $0.0447 | +290% |
| Gemini 3.8 Flash (high) | $0.0098 | $0.0302 | +207% |
| DeepSeek V4 Pro (max) | $0.0050 | $0.0186 | +273% |
Read the right-hand column and the problem is obvious. Over the two rewrites GLM-5.3 got 3.9 times more expensive per point and Claude Fable 5.1 got 2.5 times more expensive per point. Astra is on a shorter clock, having only existed for one of the two steps, and moved 1.3 times over that step against 1.6 times for Fable 5.1 on the same step. Whichever pair you take, the change is not a uniform recalibration you could divide out and carry on. Any cost-per-point table computed before September 4 is wrong by a different factor in every row, which is the specific way a comparison stops being a comparison.
Four models swapped places while sitting perfectly still
Rank the same eight models by cost per point under v4.1.1 and under v4.3. First, second, seventh and eighth hold. The whole middle reshuffles.
| Rank | Cheapest per point, v4.1.1 | Cheapest per point, v4.3 |
|---|---|---|
| 1 | DeepSeek V4 Pro (max) | DeepSeek V4 Pro (max) |
| 2 | Gemini 3.8 Flash (high) | Gemini 3.8 Flash (high) |
| 3 | GLM-5.3 (max) | Grok 4.6 (high) |
| 4 | Kimi K3 (max) | GPT-5.6 Sol (max) |
| 5 | Grok 4.6 (high) | GLM-5.3 (max) |
| 6 | GPT-5.6 Sol (max) | Kimi K3 (max) |
| 7 | Claude Opus 5 (max) | Claude Opus 5 (max) |
| 8 | Claude Fable 5.1 (max) | Claude Fable 5.1 (max) |
Anyone who picked GLM-5.3 in August because it was the third-cheapest frontier model per point of measured intelligence made a defensible choice on the evidence available, and would land on Grok 4.6 today from the same reasoning. Neither model moved. We would not put much weight on third against fourth here either way, since Grok 4.6 and GPT-5.6 Sol are separated by 0.9% on this metric, which is inside anybody's noise floor.
The raw index has the same problem in a smaller way. GLM-5.3-Flash was ahead of GPT-5.6 Terra under v4.1.1 and behind it under v4.2. Kimi K3 was ahead of GLM-5.3 under v4.1.1 and v4.2 and behind it under v4.3. Gemini 3.8 Flash was ahead of GPT-5.6 Terra for two versions and is behind it now. Write "Kimi K3 beats GLM-5.3" on September 5 and you were right. The same sentence is wrong today, and nothing shipped in between.
A benchmark and the harness that runs it were replaced in the same step
Ten percent of the index is Terminal-Bench. In v4.3 three things about that slice changed simultaneously: the task set went from 89 tasks to 66, the agent harness went from Terminus 2 to mini-SWE-agent v2.4.6, and the budget went from 250 episodes to 500 steps. Three confounded variables in one slot, and the harness is the one that bites.
Snorkel, which hosts a Terminal-Bench 4.0 leaderboard and funds the benchmark through a grants programme, mostly scores each model with the vendor's own agent: Claude Code for Claude, Codex for OpenAI, Grok Build for xAI. Artificial Analysis runs everything through mini-SWE-agent instead. Same benchmark version, same models, same week, two harnesses. Terminal-Bench itself is hosted by Harbor and the Laude Institute rather than by either of them.
| Model | AA, mini-SWE-agent | Snorkel, native agent | Gap |
|---|---|---|---|
| GPT-6 Astra (max) | 59.1% | 58.2% ± 2.8 | +0.9 |
| Claude Fable 5.1 (max) | 52.0% | 57.9% ± 3.8 | -5.9 |
| Claude Opus 5 (max) | 49.0% | 51.8% ± 3.4 | -2.8 |
| GPT-5.6 Sol (max) | 39.9% | 37.3% ± 3.8 | +2.6 |
Under Artificial Analysis' harness, Astra beats Fable 5.1 on Terminal-Bench by 7.1 points. Under Snorkel's, it beats it by 0.3, which is inside both error bars. Fable 5.1 is the model with the largest disagreement between the two, at 5.9 points, and that exceeds Snorkel's stated band of plus or minus 3.8. The most economical explanation is that Claude Code is a scaffold built around Claude and mini-SWE-agent deliberately is not built around anything. Neither choice is wrong. They just are not the same measurement, and the v4.3 narrative that Astra leads on coding rests on picking one of them.
A correction to our own August 17 post while we are here. That piece used Terminal-Bench 3.0 figures from Snorkel's leaderboard. Artificial Analysis never used 3.0 at all: its path went Terminal-Bench Hard, then v2.1, then straight to v4.0. So those numbers were never index-aligned, and neither Astra nor Fable 5.1 has a 3.0 score to compare against. Claude Opus 5 is the only bridge we have, and it reads 42.7% on Terminal-Bench 3.0, 49.0% on Artificial Analysis' 4.0, and 51.8% on Snorkel's 4.0. Three numbers, one model, one month, no model change. Only the last step there is a harness effect: the 3.0 run and Artificial Analysis' 4.0 run both use mini-SWE-agent, so 42.7 to 49.0 is the task set changing and 49.0 to 51.8 is the scaffold changing.
The warning exists. It is attached to everything except the number people quote.
Artificial Analysis knows the failure mode and writes about it carefully at the level of individual evaluations. Its methodology page says of AA-LCR that the new version "adds a system prompt to clarify grading instructions, corrects 16 answer keys" and then, flatly: "Scores are not directly comparable with v1.0." It says its GDP.pdf implementation and Surge's "are not directly comparable" because the document input and judge differ. It says its EnterpriseOps-Gym-AA numbers "are not directly comparable to results reported in the paper" because it uses its own harness. And it says its Harvey LAB-AA reimplementation is "not directly comparable to Harvey's own published results".
Four times, precisely, for four constituent benchmarks, plus a fifth near-equivalent warning discouraging direct comparison on HLE. We searched the rendered text of the v4.2 article, the v4.3 article, the methodology page and the evaluations page for any equivalent statement about the Intelligence Index itself and there is none. No comparability note, no conversion table, no restatement of prior versions, no archive of old leaderboards. The composite that gets quoted in headlines is the one number carrying no version warning, and it is the one that moved 12 to 18 points in four days.
One inconsistency is visible on the site right now. The Intelligence Index evaluation page, whose own title reads "Artificial Analysis Intelligence Index v4.3", answers its FAQ with "Four categories each contribute 25%: agents, coding, general capability, and scientific reasoning", while the methodology page and the v4.3 article both say 30% agents, 20% coding, 30% general, 20% scientific. Those are the v4.0 weights sitting live under a v4.3 header. The v4.2 announcement, for its part, carries no superseded notice, though a "Read the latest" module at the foot of it does lead with v4.3.
One of the stalest pages on the internet belongs to Artificial Analysis
On July 17 Artificial Analysis published an article whose URL slug reads, in part, "six labs now field a model above 50 on the artificial analysis intelligence index". Under v4.3 exactly three models clear 50: Claude Fable 5.1 at 53.37, GPT-6 Astra at 52.81 and Claude Opus 5 at 50.70. That is two labs, not six. The article is live, the claim is in the URL, and there is no notice on it.
The same pattern holds for the launch write-ups. The September 1 Fable 5.1 article still says it scores 66 and costs "$3.76 per Intelligence Index task", and that it is 1.6 times Claude Opus 5 on cost. Today those read 53.37, $7.63, and 1.30 times. The September 3 Astra article still says 61 and level with Sol. None of them are annotated. Old articles keep old numbers and the leaderboard silently reflows underneath them, which is a reasonable engineering decision and a rough deal for anyone who cited the article.
Who is serving the old numbers today, this site included
We checked the leaderboards that redistribute this index. Two of the four stamp a version on the number, and the more careful of those two is still a revision behind.
| Site | Version stamped? | What it currently shows | Status |
|---|---|---|---|
| whatllm.org | Yes, v4.2 | Fable 5.1 at 56.8, best task cost $6.12 | 1 revision behind |
| collectivebrain.de | No | Opus 5 at 63, Fable 5 at 62, cost per task in euros | 2 revisions behind |
| benchlm.ai | No | Fable 5.1 at 65.7%, rendered as a percentage | 2 revisions behind |
| vertu.com | Yes, v4.0 | Gemini 3.1 Pro at 57, Opus 4.6 at 53 | 4 revisions behind |
| tokencost.app/rankings | No | 104 models, fetched 2026-07-27 | 3 revisions behind |
The last row is ours. Our rankings pages carry an Artificial Analysis Intelligence Index for 104 models, fetched on July 27 from the v2 API, which puts them under v4.1 and three revisions behind the live board. The file records the date it was fetched, which is more than most, and records nothing at all about which index version produced the numbers, which is the same gap we are describing in everybody else. We have added a version field and a refresh to the queue behind this post.
Note also that whatllm.org is the site doing this best and it still got caught. It stamps "Index v4.2" and timestamps its snapshot to 15:56 UTC on September 7, hours after v4.3 went live. Version-stamping does not save you on its own. It just means the reader can tell.
The half of this that does not rot
None of the above is an argument against the index. Adding private held-out sets, cutting a saturated benchmark and moving toward agentic tasks all make it a better instrument, and Artificial Analysis is one of very few outfits publishing what a benchmark run costs at all. The problem is downstream of the index and it is a provenance problem, not a quality one.
Three things worth doing, in order of how much they save you. First, never write down an index score without the version and the date beside it. A bare 53 is unusable in six weeks and there is no way to recover what it meant. Second, treat any cost-per-point figure as valid only within one index version, because the deflator is not constant and comparisons across versions reorder. Third, if the decision is worth real money, the per-token rate card is the durable half of it. $10.00 and $50.00 has been true for Fable 5.1 and Astra all week while everything else on this page moved.
And if you want the comparison that does not rot, run your own workload. Our calculator prices your token counts against the current cards, and the pricing table is dated per row. A rate card is a fact about a contract. An index score is a fact about a methodology, and this month the methodology changed twice.
Objections, and what we said back
Why did every Intelligence Index score drop in September 2026?
Because the index was rewritten twice, on September 4 and September 7. Between them the two versions added two benchmarks worth 25% of total weight, deleted GPQA Diamond, replaced τ³-Banking with AutomationBench-AA, and swapped Terminal-Bench v2.1 for v4.0 on a different agent harness. No model was retrained. Claude Fable 5.1 went 65.65, then 56.76, then 53.37, and every model on the board before September 4 fell between 12.3 and 18.1 points.
Can I convert an old score to the new scale?
No. The v4.2 to v4.3 step alone multiplied scores by between 0.74 and 0.97 depending on the model, so there is no single deflator to divide out. Artificial Analysis has published no conversion, and it keeps no archive of prior-version leaderboards, so recovering the old numbers means the Wayback Machine. Cost per point rose between 154% and 290% across the models measured through both rewrites, and rankings genuinely reordered.
Did GPT-6 Astra actually catch Claude Fable 5.1?
On the current methodology they are 0.560 points apart, which rounds to a tie at 53. Three days earlier, on v4.2, they were 2.101 points apart with Fable 5.1 ahead. Astra also happened to lose less than any other model on the board in that step, keeping 96.6% of its score against Fable 5.1's 94.0%. Nothing about either model changed. On cost the gap is real and large either way: $3.26 per task against $7.63, on identical $10.00 and $50.00 rate cards, because Astra emits far fewer tokens getting there.
Does Artificial Analysis warn that versions are not comparable?
For individual benchmarks, yes, and precisely. The phrase "not directly comparable" appears four times on the methodology page: for AA-LCR v1.0 against v1.1, for its GDP.pdf implementation against Surge's, for EnterpriseOps-Gym-AA against the paper, and for Harvey LAB-AA against Harvey's own results. For the composite Intelligence Index there is no such statement anywhere we could find, which is the one place it would do the most good.
Is Terminal-Bench 4.0 measuring what 2.1 measured?
Not really, and three things changed at once: 89 tasks to 66, Terminus 2 to mini-SWE-agent v2.4.6, and 250 episodes to 500 steps. The harness is measurable on its own. On the same 4.0 task set, Artificial Analysis scores Claude Fable 5.1 at 52.0% with mini-SWE-agent while Snorkel scores it 57.9% with Claude Code, a 5.9 point gap that is wider than Snorkel's own plus or minus 3.8 error bar.
Provenance, including the places ours is weaker
Index scores, cost per task and output-token counts are the values Artificial Analysis embeds as structured data on its own leaderboard and model pages: live for v4.3, and via Wayback captures of the same URL dated September 3 and September 7 for v4.1.1 and v4.2. Weights, version history and the comparability language come from the methodology page and the v4.3 and v4.2 announcements. Terminal-Bench 4.0 scores and error bars are from Snorkel's leaderboard.
Rate cards are vendor-primary where we could reach them: Anthropic, Z.AI and Google. OpenAI's pricing page returned 403 to every automated fetch we tried, so the GPT-6 Astra, Sol, Terra and Luna figures here are cross-checked against Artificial Analysis' own pricing dataset rather than read off openai.com. Two independent sources agree on every row, but they are not the vendor's page and we would rather say so.
Three things we could not pin down. We could not verify from first-party vendor pages that labs cite these index scores in launch marketing, because anthropic.com/news returned 404 and openai.com returned 403; the citations we can evidence are press and analyst write-ups, which is a weaker claim than the one we set out to make. Artificial Analysis names no defect in τ³-Banking to explain its removal, only a case for what replaced it, so that gap is reported as an absence and not interpreted. And Astra's 61 is quoted from the September 3 article, whose chart is captioned v4.1.1, rather than recovered from a leaderboard capture, because Astra had not yet appeared on the board when that capture was taken.
Price the work, not the ruler
Every rate card in this post is on the pricing table with the date we last checked it against the vendor's own page. If you want a number that will still mean the same thing in November, start there and price your own tokens.