LLM Rankings: Every List, With the Method
9 ranked lists over 144 models you can actually buy today. Every score traces to a benchmark a provider or an independent lab published. Models with no published score are listed separately, never ranked at the bottom on a zero they never earned.
Ranked on quality among the 57 models with a published SWE-bench Verified or Terminal-Bench 2.1 score (7 have SWE-bench Verified, 55 have Terminal-Bench 2.1, 5 have both). A further 87 models are listed on price and context only, because they have no published score and a missing score is not a zero. 6 models are excluded from the ranking entirely.
Pricing as of July 2026. Nothing here is estimated, interpolated or inferred: if a model has no published score for a benchmark, that cell is empty and the model is scored only on what it does have. 33 of the scored models were benchmarked only at a raised reasoning effort rather than the default a plain API call gives you; those rows are marked non-default tier and every one of them is named below, because a max-effort score next to a medium-effort one is not a like-for-like comparison and costs more output tokens than the price on the row assumes. Read the full method.
Every ranking
One list, one page, one question. Each row states what the list sorts on, how many models are in it, and who currently holds first place.
| List | Sorted on | Models | Currently #1 |
|---|---|---|---|
| The Best LLMs by Benchmark ScoreWhich model scores highest on the published coding benchmarks? | Quality score | 57 | Claude Opus 4.887.0 |
| The Best Value LLMs per DollarWhich model buys the most benchmark quality per dollar? | Quality per $/1M | 18 | Nex-N2-Pro75.60 |
| The Best LLMs by Context WindowWhich model holds the most tokens in one request? | Context window | 144 | Grok 4.1 Fast2M |
| The Cheapest LLMs With a 1M-Token Context WindowWhich million-token model costs least to run? | Blended $/1M | 50 | Llama 4 Scout$0.135 |
| The Best LLMs on Terminal-Bench 2.1Which model solves the most Terminal-Bench 2.1 tasks? | Terminal-Bench 2.1 | 55 | GPT-5.6 Sol86.1 |
| The Best LLMs on Terminal-Bench HardWhich model holds up on the hard Terminal-Bench subset? | Terminal-Bench Hard | 86 | GPT-5.6 Sol62.9 |
| The Best LLMs on SWE-bench VerifiedWhich model solves the most SWE-bench Verified issues? | SWE-bench Verified | 7 | Claude Opus 4.888.6 |
| The Best LLMs on GPQA DiamondWhich model answers the most graduate-level science questions? | GPQA Diamond | 103 | Kimi K393.5 |
| The Best LLMs for Tool Use on tau2-benchWhich model is most reliable at calling tools? | tau2-bench | 89 | GLM-5.299.1 |
| Intelligence indexThe leaderboard already ranks these models on the same index, live and sortable, with speed and price alongside. A second static copy would compete with it for the same query and answer it worse. | AA Intelligence Index | 105 | Claude Fable 559.9 |
By tool
The lists above rank an open field on a published number. These six pages answer a different question — “I use this tool, what should I pick?” — on an editorial roster, with each page weighting the suites for its own workload and writing up every model it recommends.
The method
A ranking is only worth citing if you can check it. Everything the formula uses is below: the weights, what data each model actually has, how many models qualified, and which models were thrown out and why. Given a model's published scores and its published prices, you can reproduce every number on every list with a calculator.
1. The quality score
The quality score is a weighted mean of the published pass-rate benchmarks a model has. Both suites report the same kind of number, the percentage of a fixed task set that the model solved, so they can be averaged without rescaling anything.
| Benchmark | Weight | Unit | Models with it | Why this weight |
|---|---|---|---|---|
| SWE-bench Verified | 0.60 | % of tasks solved (pass@1) | 7 | The most widely reported and most often independently replicated agentic coding suite, so it carries the larger share. |
| Terminal-Bench 2.1 | 0.40 | % of tasks solved | 55 | Broader agentic/terminal work, but newer and more often vendor-reported, so it carries the smaller share. |
Why the weights are renormalised
Most models have published only one of the two suites. If the missing suite were treated as a 0, a model scoring 80 on SWE-bench alone would come out at 48, below models that are plainly worse. So the divisor is the sum of the weights actually present: a SWE-bench-only model scoring 80 scores 80. The Evidence column says how many suites went into each score, because a score from one suite is a weaker claim than a score from two, and you should be able to see which you are looking at. Of the 57 ranked models, 5 published both suites.
What is deliberately left out: Artificial Analysis Intelligence Index
An index on its own scale, not a percentage of tasks solved. Rescaling it onto the pass-rate scale would mean inventing a ceiling, so it is reported separately rather than blended. 105 models carry it. They are ranked on it, live and sortable, on the leaderboard, separately from the quality score and never blended into it.
2. The price, and quality per dollar
value = quality / blended $/1M
Two published prices are collapsed into one comparable figure using a 3:1 input-to-output tokens workload mix, because agentic and coding traffic reads far more than it writes. That mix is a modelling assumption, not a measurement, and it moves the value ranking, so it is stated here rather than buried. Our cheapest models page uses a simpler 1:1 average over the whole catalogue with no quality floor, which answers a different question.
Models priced at $0 (open weights with no metered API) get no value score: you cannot divide by a price that does not exist. The best value list uses a floor of 70 quality points. That floor is an editorial line, not a measurement, and it is the only thing that separates that list from a plain price list.
3. Coverage: what data each model has
| Models in the catalogue | 150 |
| Eligible for ranking | 144 |
| Excluded before ranking | 6 |
| Ranked on quality (has at least one pass-rate suite) | 57 |
| With SWE-bench Verified | 7 |
| With Terminal-Bench 2.1 | 55 |
| With both suites | 5 |
| With an Artificial Analysis index (ranked separately) | 105 |
| No published score: listed on price and context only | 87 |
The 87 models with no published pass-rate score are not ranked below the scored ones, and they are not given a score of zero. They are absent from the quality lists entirely, and they compete on equal terms in longest context, which needs no benchmark data. Browse them on the model directory or the pricing table. An absent number is a fact about what the provider published, not a gap for us to fill.
4. Exclusions: 6 models, and why
A ranking is a buying recommendation. Putting a model at #2 asserts that you can go and pay that price today, so models you cannot buy are removed rather than footnoted: 2 partner-gated, 3 retired, and 1 whose price we could not read off a provider's own page. Their model pages stay live.
| Model | Reason | Why it disqualifies |
|---|---|---|
| Claude Mythos 5Anthropic | partner-gated | Access is gated to vetted partners or a waitlist, so a normal reader cannot buy it at any price. |
| Claude Mythos PreviewAnthropic | partner-gated, price-unverified | Access is gated to vetted partners or a waitlist, so a normal reader cannot buy it at any price. The provider does not publish a rate for this model, so the figure here comes from somewhere weaker than a provider page. |
| Claude Sonnet 4Anthropic | retired | Shut down by the provider. The price is history, not an option. |
| Gemini 2.0 FlashGoogle | retired | Shut down by the provider. The price is history, not an option. |
| Gemini 2.0 Flash-LiteGoogle | retired | Shut down by the provider. The price is history, not an option. |
| Hunyuan HY3 PreviewTencent | price-unverified | The provider does not publish a rate for this model, so the figure here comes from somewhere weaker than a provider page. |
The formula is deterministic and pure: no randomness, no date that shifts the order. The only time-dependent input is a provider's own announced price change on its own announced date, which corrects the price rather than drifting it. Two builds of the same catalogue produce the same ranking.
5. Measured off the default tier: 33 models, named
Artificial Analysis lists reasoning-effort variants as separate entries and does not benchmark every model at the tier a plain API call gives you. Where it did not, the score below was measured at a higher effort than the price on the same row assumes, so the row is not directly comparable to one measured at the default and will burn more output tokens than its price implies. An unmarked mixed-tier table quietly favours whichever model was tested hardest, so every affected model is named here with the exact Artificial Analysis entry used. 33 of the 105 indexed models are affected.
| Model | Artificial Analysis entry used | AA index | Terminal-Bench |
|---|---|---|---|
| Qwen 3.5 27BAlibaba | Qwen3.5 27B (Reasoning) | 33.8 | not published |
| Qwen3.5-9BAlibaba | Qwen3.5 9B (Reasoning) | 21.4 | 29.2 |
| Claude 3.7 SonnetAnthropic | Claude 3.7 Sonnet (Reasoning) | 27.1 | not published |
| Claude Opus 4.5Anthropic | Claude Opus 4.5 (Reasoning) | 40.8 | not published |
| Claude Opus 4.6Anthropic | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) | 43.7 | not published |
| Claude Opus 4.7Anthropic | Claude Opus 4.7 (Adaptive Reasoning, Max Effort) | 53.5 | 83.1 |
| Claude Sonnet 4.6Anthropic | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) | 47.2 | 71.2 |
| Claude Sonnet 5Anthropic | Claude Sonnet 5 (Adaptive Reasoning, Max Effort) | 53.4 | 80.5 |
| DeepSeek V4-FlashDeepSeek | DeepSeek V4 Flash (Reasoning, High Effort) | 37.5 | 56.9 |
| DeepSeek V4-ProDeepSeek | DeepSeek V4 Pro (Reasoning, High Effort) | 43.1 | 64.8 |
| Gemini 2.5 FlashGoogle | Gemini 2.5 Flash (Reasoning) | 20.1 | not published |
| Gemini 2.5 Flash-LiteGoogle | Gemini 2.5 Flash-Lite (Reasoning) | 11.4 | not published |
| Gemini 3 Flash ReasoningGoogle | Gemini 3 Flash Preview (Reasoning) | 37.8 | not published |
| Gemma 4 12BGoogle | Gemma 4 12B (Reasoning) | 21.8 | 27.3 |
| Gemma 4 26B A4BGoogle | Gemma 4 26B A4B (Reasoning) | 25.7 | 39.0 |
| Gemma 4 31BGoogle | Gemma 4 31B (Reasoning) | 29.4 | 43.4 |
| Devstral SmallMistral | Devstral Small (Jul '25) | 9.3 | not published |
| Mistral Small 3.2Mistral | Mistral Small (Sep '24) | 4.7 | not published |
| Mistral Small 4Mistral | Mistral Small 4 (Reasoning) | 19.6 | 21.0 |
| Kimi K2.5Moonshot | Kimi K2.5 (Reasoning) | 35.4 | 45.7 |
| GPT-4oOpenAI | GPT-4o (Nov '24) | 11.2 | not published |
| GPT-5.1OpenAI | GPT-5.1 (high) | 36.9 | 52.4 |
| GPT-5.4OpenAI | GPT-5.4 (low) | 39.1 | not published |
| GPT-OSS 120BOpenAI | gpt-oss-120b (low) | 14.9 | 13.9 |
| GPT-OSS 20BOpenAI | gpt-oss-20b (low) | 14.3 | not published |
| Grok 4.1 FastxAI | Grok 4.1 Fast (Reasoning) | 30.6 | not published |
| Grok 4.1 Fast ReasoningxAI | Grok 4.1 Fast (Reasoning) | 30.6 | not published |
| Grok 4.20xAI | Grok 4.20 0309 v2 (Reasoning) | 37.0 | not published |
| GLM-4.7Zhipu | GLM-4.7 (Reasoning) | 33.7 | 45.3 |
| GLM-4.7-flashZhipu | GLM-4.7-Flash (Reasoning) | 22.9 | not published |
| GLM-5Zhipu | GLM-5 (Reasoning) | 39.5 | not published |
| GLM-5.1Zhipu | GLM-5.1 (Reasoning) | 40.2 | 61.8 |
| GLM-5.2Zhipu | GLM-5.2 (max) | 51.1 | 77.9 |
6. Which suites each provider actually publishes
Every count below is of models whose score we could read from a published source. It is the answer to why the SWE-bench Verified list holds 7 models and the GPQA list holds far more: providers pick which suites to report, and most report neither of the two the quality score uses.
| Provider | Models | SWE-bench V. | Terminal-Bench | TB Hard | GPQA | tau2 | AA index |
|---|---|---|---|---|---|---|---|
| OpenAI | 33 | 0 | 9 | 22 | 26 | 23 | 27 |
| 17 | 0 | 8 | 9 | 11 | 9 | 11 | |
| Mistral | 15 | 1 | 5 | 8 | 11 | 9 | 11 |
| xAI | 13 | 0 | 1 | 6 | 8 | 6 | 8 |
| Alibaba | 12 | 1 | 5 | 9 | 9 | 9 | 9 |
| Anthropic | 12 | 2 | 5 | 7 | 8 | 7 | 8 |
| Moonshot | 6 | 0 | 4 | 4 | 5 | 4 | 5 |
| Zhipu | 6 | 0 | 3 | 6 | 6 | 6 | 6 |
| DeepSeek | 5 | 0 | 2 | 3 | 3 | 3 | 3 |
| Meta | 4 | 0 | 3 | 2 | 3 | 2 | 3 |
| MiniMax | 3 | 0 | 2 | 3 | 3 | 3 | 3 |
| Amazon | 2 | 0 | 0 | 1 | 1 | 1 | 1 |
| ByteDance | 2 | 0 | 0 | 0 | 0 | 0 | 0 |
| Cohere | 2 | 0 | 2 | 2 | 2 | 2 | 2 |
| Kwaipilot | 2 | 0 | 1 | 0 | 0 | 0 | 0 |
| NVIDIA | 2 | 1 | 0 | 0 | 0 | 0 | 1 |
| Xiaomi | 2 | 1 | 1 | 2 | 2 | 2 | 2 |
| Inception Labs | 1 | 0 | 1 | 1 | 1 | 1 | 1 |
| Meituan | 1 | 0 | 1 | 0 | 1 | 0 | 1 |
| Microsoft | 1 | 0 | 0 | 0 | 0 | 0 | 0 |
| Nex AGI | 1 | 1 | 1 | 0 | 1 | 1 | 1 |
| Perplexity | 1 | 0 | 0 | 0 | 1 | 0 | 1 |
| StepFun | 1 | 0 | 1 | 1 | 1 | 1 | 1 |
Questions about the method
Every answer below is a lookup against the same data the lists are built from.
LLM Rankings: Every List, With the Method. Pricing as of July 2026.