Skip to main content
TokenCost logoTokenCost

LLM Rankings: Every List, With the Method

9 ranked lists over 144 models you can actually buy today. Every score traces to a benchmark a provider or an independent lab published. Models with no published score are listed separately, never ranked at the bottom on a zero they never earned.

Ranked on quality among the 57 models with a published SWE-bench Verified or Terminal-Bench 2.1 score (7 have SWE-bench Verified, 55 have Terminal-Bench 2.1, 5 have both). A further 87 models are listed on price and context only, because they have no published score and a missing score is not a zero. 6 models are excluded from the ranking entirely.

Pricing as of July 2026. Nothing here is estimated, interpolated or inferred: if a model has no published score for a benchmark, that cell is empty and the model is scored only on what it does have. 33 of the scored models were benchmarked only at a raised reasoning effort rather than the default a plain API call gives you; those rows are marked non-default tier and every one of them is named below, because a max-effort score next to a medium-effort one is not a like-for-like comparison and costs more output tokens than the price on the row assumes. Read the full method.

Every ranking

One list, one page, one question. Each row states what the list sorts on, how many models are in it, and who currently holds first place.

ListSorted onModelsCurrently #1
The Best LLMs by Benchmark ScoreWhich model scores highest on the published coding benchmarks?Quality score57Claude Opus 4.887.0
The Best Value LLMs per DollarWhich model buys the most benchmark quality per dollar?Quality per $/1M18Nex-N2-Pro75.60
The Best LLMs by Context WindowWhich model holds the most tokens in one request?Context window144Grok 4.1 Fast2M
The Cheapest LLMs With a 1M-Token Context WindowWhich million-token model costs least to run?Blended $/1M50Llama 4 Scout$0.135
The Best LLMs on Terminal-Bench 2.1Which model solves the most Terminal-Bench 2.1 tasks?Terminal-Bench 2.155GPT-5.6 Sol86.1
The Best LLMs on Terminal-Bench HardWhich model holds up on the hard Terminal-Bench subset?Terminal-Bench Hard86GPT-5.6 Sol62.9
The Best LLMs on SWE-bench VerifiedWhich model solves the most SWE-bench Verified issues?SWE-bench Verified7Claude Opus 4.888.6
The Best LLMs on GPQA DiamondWhich model answers the most graduate-level science questions?GPQA Diamond103Kimi K393.5
The Best LLMs for Tool Use on tau2-benchWhich model is most reliable at calling tools?tau2-bench89GLM-5.299.1
Intelligence indexThe leaderboard already ranks these models on the same index, live and sortable, with speed and price alongside. A second static copy would compete with it for the same query and answer it worse.AA Intelligence Index105Claude Fable 559.9

By tool

The lists above rank an open field on a published number. These six pages answer a different question — “I use this tool, what should I pick?” — on an editorial roster, with each page weighting the suites for its own workload and writing up every model it recommends.

The method

A ranking is only worth citing if you can check it. Everything the formula uses is below: the weights, what data each model actually has, how many models qualified, and which models were thrown out and why. Given a model's published scores and its published prices, you can reproduce every number on every list with a calculator.

1. The quality score

The quality score is a weighted mean of the published pass-rate benchmarks a model has. Both suites report the same kind of number, the percentage of a fixed task set that the model solved, so they can be averaged without rescaling anything.

quality = (0.6 x SWE-bench Verified + 0.4 x Terminal-Bench 2.1) / (sum of the weights of the suites present)
BenchmarkWeightUnitModels with itWhy this weight
SWE-bench Verified0.60% of tasks solved (pass@1)7The most widely reported and most often independently replicated agentic coding suite, so it carries the larger share.
Terminal-Bench 2.10.40% of tasks solved55Broader agentic/terminal work, but newer and more often vendor-reported, so it carries the smaller share.

Why the weights are renormalised

Most models have published only one of the two suites. If the missing suite were treated as a 0, a model scoring 80 on SWE-bench alone would come out at 48, below models that are plainly worse. So the divisor is the sum of the weights actually present: a SWE-bench-only model scoring 80 scores 80. The Evidence column says how many suites went into each score, because a score from one suite is a weaker claim than a score from two, and you should be able to see which you are looking at. Of the 57 ranked models, 5 published both suites.

What is deliberately left out: Artificial Analysis Intelligence Index

An index on its own scale, not a percentage of tasks solved. Rescaling it onto the pass-rate scale would mean inventing a ceiling, so it is reported separately rather than blended. 105 models carry it. They are ranked on it, live and sortable, on the leaderboard, separately from the quality score and never blended into it.

2. The price, and quality per dollar

blended $/1M = 0.75 x input + 0.25 x output
value = quality / blended $/1M

Two published prices are collapsed into one comparable figure using a 3:1 input-to-output tokens workload mix, because agentic and coding traffic reads far more than it writes. That mix is a modelling assumption, not a measurement, and it moves the value ranking, so it is stated here rather than buried. Our cheapest models page uses a simpler 1:1 average over the whole catalogue with no quality floor, which answers a different question.

Models priced at $0 (open weights with no metered API) get no value score: you cannot divide by a price that does not exist. The best value list uses a floor of 70 quality points. That floor is an editorial line, not a measurement, and it is the only thing that separates that list from a plain price list.

3. Coverage: what data each model has

Models in the catalogue150
Eligible for ranking144
Excluded before ranking6
Ranked on quality (has at least one pass-rate suite)57
With SWE-bench Verified7
With Terminal-Bench 2.155
With both suites5
With an Artificial Analysis index (ranked separately)105
No published score: listed on price and context only87

The 87 models with no published pass-rate score are not ranked below the scored ones, and they are not given a score of zero. They are absent from the quality lists entirely, and they compete on equal terms in longest context, which needs no benchmark data. Browse them on the model directory or the pricing table. An absent number is a fact about what the provider published, not a gap for us to fill.

4. Exclusions: 6 models, and why

A ranking is a buying recommendation. Putting a model at #2 asserts that you can go and pay that price today, so models you cannot buy are removed rather than footnoted: 2 partner-gated, 3 retired, and 1 whose price we could not read off a provider's own page. Their model pages stay live.

ModelReasonWhy it disqualifies
Claude Mythos 5Anthropicpartner-gatedAccess is gated to vetted partners or a waitlist, so a normal reader cannot buy it at any price.
Claude Mythos PreviewAnthropicpartner-gated, price-unverifiedAccess is gated to vetted partners or a waitlist, so a normal reader cannot buy it at any price. The provider does not publish a rate for this model, so the figure here comes from somewhere weaker than a provider page.
Claude Sonnet 4AnthropicretiredShut down by the provider. The price is history, not an option.
Gemini 2.0 FlashGoogleretiredShut down by the provider. The price is history, not an option.
Gemini 2.0 Flash-LiteGoogleretiredShut down by the provider. The price is history, not an option.
Hunyuan HY3 PreviewTencentprice-unverifiedThe provider does not publish a rate for this model, so the figure here comes from somewhere weaker than a provider page.

The formula is deterministic and pure: no randomness, no date that shifts the order. The only time-dependent input is a provider's own announced price change on its own announced date, which corrects the price rather than drifting it. Two builds of the same catalogue produce the same ranking.

5. Measured off the default tier: 33 models, named

Artificial Analysis lists reasoning-effort variants as separate entries and does not benchmark every model at the tier a plain API call gives you. Where it did not, the score below was measured at a higher effort than the price on the same row assumes, so the row is not directly comparable to one measured at the default and will burn more output tokens than its price implies. An unmarked mixed-tier table quietly favours whichever model was tested hardest, so every affected model is named here with the exact Artificial Analysis entry used. 33 of the 105 indexed models are affected.

ModelArtificial Analysis entry usedAA indexTerminal-Bench
Qwen 3.5 27BAlibabaQwen3.5 27B (Reasoning)33.8not published
Qwen3.5-9BAlibabaQwen3.5 9B (Reasoning)21.429.2
Claude 3.7 SonnetAnthropicClaude 3.7 Sonnet (Reasoning)27.1not published
Claude Opus 4.5AnthropicClaude Opus 4.5 (Reasoning)40.8not published
Claude Opus 4.6AnthropicClaude Opus 4.6 (Adaptive Reasoning, Max Effort)43.7not published
Claude Opus 4.7AnthropicClaude Opus 4.7 (Adaptive Reasoning, Max Effort)53.583.1
Claude Sonnet 4.6AnthropicClaude Sonnet 4.6 (Adaptive Reasoning, Max Effort)47.271.2
Claude Sonnet 5AnthropicClaude Sonnet 5 (Adaptive Reasoning, Max Effort)53.480.5
DeepSeek V4-FlashDeepSeekDeepSeek V4 Flash (Reasoning, High Effort)37.556.9
DeepSeek V4-ProDeepSeekDeepSeek V4 Pro (Reasoning, High Effort)43.164.8
Gemini 2.5 FlashGoogleGemini 2.5 Flash (Reasoning)20.1not published
Gemini 2.5 Flash-LiteGoogleGemini 2.5 Flash-Lite (Reasoning)11.4not published
Gemini 3 Flash ReasoningGoogleGemini 3 Flash Preview (Reasoning)37.8not published
Gemma 4 12BGoogleGemma 4 12B (Reasoning)21.827.3
Gemma 4 26B A4BGoogleGemma 4 26B A4B (Reasoning)25.739.0
Gemma 4 31BGoogleGemma 4 31B (Reasoning)29.443.4
Devstral SmallMistralDevstral Small (Jul '25)9.3not published
Mistral Small 3.2MistralMistral Small (Sep '24)4.7not published
Mistral Small 4MistralMistral Small 4 (Reasoning)19.621.0
Kimi K2.5MoonshotKimi K2.5 (Reasoning)35.445.7
GPT-4oOpenAIGPT-4o (Nov '24)11.2not published
GPT-5.1OpenAIGPT-5.1 (high)36.952.4
GPT-5.4OpenAIGPT-5.4 (low)39.1not published
GPT-OSS 120BOpenAIgpt-oss-120b (low)14.913.9
GPT-OSS 20BOpenAIgpt-oss-20b (low)14.3not published
Grok 4.1 FastxAIGrok 4.1 Fast (Reasoning)30.6not published
Grok 4.1 Fast ReasoningxAIGrok 4.1 Fast (Reasoning)30.6not published
Grok 4.20xAIGrok 4.20 0309 v2 (Reasoning)37.0not published
GLM-4.7ZhipuGLM-4.7 (Reasoning)33.745.3
GLM-4.7-flashZhipuGLM-4.7-Flash (Reasoning)22.9not published
GLM-5ZhipuGLM-5 (Reasoning)39.5not published
GLM-5.1ZhipuGLM-5.1 (Reasoning)40.261.8
GLM-5.2ZhipuGLM-5.2 (max)51.177.9

6. Which suites each provider actually publishes

Every count below is of models whose score we could read from a published source. It is the answer to why the SWE-bench Verified list holds 7 models and the GPQA list holds far more: providers pick which suites to report, and most report neither of the two the quality score uses.

ProviderModelsSWE-bench V.Terminal-BenchTB HardGPQAtau2AA index
OpenAI330922262327
Google1708911911
Mistral1515811911
xAI13016868
Alibaba12159999
Anthropic12257878
Moonshot6044545
Zhipu6036666
DeepSeek5023333
Meta4032323
MiniMax3023333
Amazon2001111
ByteDance2000000
Cohere2022222
Kwaipilot2010000
NVIDIA2100001
Xiaomi2112222
Inception Labs1011111
Meituan1010101
Microsoft1000000
Nex AGI1110111
Perplexity1000101
StepFun1011111

Questions about the method

Every answer below is a lookup against the same data the lists are built from.

LLM Rankings: Every List, With the Method. Pricing as of July 2026.