The Best LLMs by Benchmark Score
Which model scores highest on the published coding benchmarks?
57 of the 144 buyable models carry a published pass-rate score. The other 87 are absent from this list rather than ranked at the bottom on a zero they never earned.
Method: Weighted mean of the published pass-rate suites the model has (SWE-bench Verified 0.6, Terminal-Bench 2.1 0.4), renormalised over the suites present. An exact tie breaks first to the model with more suites published, since a score corroborated by two suites is a stronger claim than the same score from one; then to the lower blended price, then model id. Pricing as of July 2026. 16 rows below carry a score measured at a non-default reasoning tier and are marked as such; every affected model is named on the hub. Read the full method.
| Rank | Model | Quality | SWE-bench V. | Terminal-Bench | Evidence | Reported by | Blended $/1M | Per $ |
|---|---|---|---|---|---|---|---|---|
| #1 | Claude Opus 4.8Anthropic | 87.0 | 88.6 | 84.6 | Both suites | Unstated / mixed | $10.00 | 8.70 |
| #2 | GPT-5.6 SolOpenAI | 86.1 | not published | 86.1 | Terminal-Bench only | Independent | $11.25 | 7.65 |
| #3 | Kimi K3Moonshot | 85.0 | not published | 85.0 | Terminal-Bench only | Independent | $6.00 | 14.17 |
| #4 | Claude Fable 5Anthropic | 84.6 | not published | 84.6 | Terminal-Bench only | Independent | $20.00 | 4.23 |
| #5 | Claude Sonnet 5Anthropic· non-default tier | 83.3 | 85.2 | 80.5 | Both suites | Unstated / mixed | $4.00 | 20.83 |
| #6 | Claude Opus 4.7Anthropic· non-default tier | 83.1 | not published | 83.1 | Terminal-Bench only | Independent | $10.00 | 8.31 |
| #7 | Grok 4.5xAI | 81.6 | not published | 81.6 | Terminal-Bench only | Independent | $3.00 | 27.20 |
| #8 | GPT-5.5OpenAI | 80.5 | not published | 80.5 | Terminal-Bench only | Independent | $11.25 | 7.16 |
| #9 | MiMo-V2-ProXiaomi | 78.0 | 78.0 | not published | SWE-bench only | Unstated / mixed | $1.50 | 52.00 |
| #10 | Muse Spark 1.1Meta | 77.9 | not published | 77.9 | Terminal-Bench only | Independent | $2.00 | 38.95 |
| #11 | GLM-5.2Zhipu· non-default tier | 77.9 | not published | 77.9 | Terminal-Bench only | Independent | $2.15 | 36.23 |
| #12 | Gemini 3.6 FlashGoogle | 77.5 | not published | 77.5 | Terminal-Bench only | Independent | $3.00 | 25.83 |
| #13 | Gemini 3.5 FlashGoogle | 76.2 | not published | 76.2 | Terminal-Bench only | Unstated / mixed | $3.38 | 22.58 |
| #14 | Nex-N2-ProNex AGI | 75.6 | 80.8 | 67.8 | Both suites | Unstated / mixed | $1.00 | 75.60 |
| #15 | Qwen3.7 MaxAlibaba | 74.5 | not published | 74.5 | Terminal-Bench only | Independent | $3.75 | 19.87 |
| #16 | GPT-5.6 TerraOpenAI | 72.3 | not published | 72.3 | Terminal-Bench only | Independent | $5.63 | 12.85 |
| #17 | Nemotron 3 Ultra 550BNVIDIA | 71.9 | 71.9 | not published | SWE-bench only | Unstated / mixed | $1.00 | 71.90 |
| #18 | Claude Sonnet 4.6Anthropic· non-default tier | 71.2 | not published | 71.2 | Terminal-Bench only | Independent | $6.00 | 11.87 |
| #19 | Kimi K2.7 CodeMoonshot | 67.4 | not published | 67.4 | Terminal-Bench only | Independent | $1.71 | 39.36 |
| #20 | Mistral Medium 3.5Mistral | 66.8 | 77.6 | 50.6 | Both suites | Unstated / mixed | $3.00 | 22.27 |
| #21 | Kimi K2.6Moonshot | 65.9 | not published | 65.9 | Terminal-Bench only | Independent | $1.71 | 38.48 |
| #22 | MiMo-V2.5-ProXiaomi | 65.2 | not published | 65.2 | Terminal-Bench only | Independent | $0.544 | 119.90 |
| #23 | MiniMax M3MiniMax | 65.2 | not published | 65.2 | Terminal-Bench only | Independent | $1.05 | 62.10 |
| #24 | DeepSeek V4-ProDeepSeek· non-default tier | 64.8 | not published | 64.8 | Terminal-Bench only | Independent | $0.544 | 119.16 |
| #25 | GLM-5.1Zhipu· non-default tier | 61.8 | not published | 61.8 | Terminal-Bench only | Independent | $2.15 | 28.74 |
| #26 | Qwen3.6-PlusAlibaba | 61.4 | not published | 61.4 | Terminal-Bench only | Independent | $0.620 | 99.06 |
| #27 | Qwen3.7 PlusAlibaba | 61.0 | not published | 61.0 | Terminal-Bench only | Independent | $0.700 | 87.14 |
| #28 | KAT-Coder-Pro v2.5Kwaipilot | 60.7 | not published | 60.7 | Terminal-Bench only | Vendor-reported | $1.29 | 46.87 |
| #29 | Qwen3 Coder NextAlibaba | 57.6 | 70.6 | 38.2 | Both suites | Unstated / mixed | $0.282 | 203.89 |
| #30 | DeepSeek V4-FlashDeepSeek· non-default tier | 56.9 | not published | 56.9 | Terminal-Bench only | Independent | $0.175 | 325.14 |
| #31 | MiniMax M2.7MiniMax | 55.4 | not published | 55.4 | Terminal-Bench only | Independent | $0.525 | 105.52 |
| #32 | Gemini 3.5 Flash-LiteGoogle | 53.6 | not published | 53.6 | Terminal-Bench only | Independent | $0.850 | 63.06 |
| #33 | GPT-5.6 LunaOpenAI | 53.2 | not published | 53.2 | Terminal-Bench only | Independent | $2.25 | 23.64 |
| #34 | GPT-5.1OpenAI· non-default tier | 52.4 | not published | 52.4 | Terminal-Bench only | Independent | $3.44 | 15.24 |
| #35 | LongCat-2.0Meituan | 50.2 | not published | 50.2 | Terminal-Bench only | Independent | $1.30 | 38.62 |
| #36 | Kimi K2.5Moonshot· non-default tier | 45.7 | not published | 45.7 | Terminal-Bench only | Independent | $1.20 | 38.08 |
| #37 | GLM-4.7Zhipu· non-default tier | 45.3 | not published | 45.3 | Terminal-Bench only | Independent | $1.00 | 45.30 |
| #38 | Gemma 4 31BGoogle· non-default tier | 43.4 | not published | 43.4 | Terminal-Bench only | Independent | $0.180 | 241.11 |
| #39 | Step 3.7 FlashStepFun | 39.3 | not published | 39.3 | Terminal-Bench only | Independent | $0.438 | 89.83 |
| #40 | Gemma 4 26B A4BGoogle· non-default tier | 39.0 | not published | 39.0 | Terminal-Bench only | Independent | $0.120 | 325.00 |
| #41 | Gemini 3.1 Flash-LiteGoogle | 31.1 | not published | 31.1 | Terminal-Bench only | Independent | $0.563 | 55.29 |
| #42 | Qwen3.5-9BAlibaba· non-default tier | 29.2 | not published | 29.2 | Terminal-Bench only | Independent | $0.075 | 389.33 |
| #43 | Gemini 2.5 ProGoogle | 28.5 | not published | 28.5 | Terminal-Bench only | Independent | $3.44 | 8.29 |
| #44 | Gemma 4 12BGoogle· non-default tier | 27.3 | not published | 27.3 | Terminal-Bench only | Independent | $0.00 | — |
| #45 | Mercury 2Inception Labs | 27.3 | not published | 27.3 | Terminal-Bench only | Independent | $0.375 | 72.80 |
| #46 | Command ACohere | 22.8 | not published | 22.8 | Terminal-Bench only | Independent | $4.38 | 5.21 |
| #47 | Command A+Cohere | 22.8 | not published | 22.8 | Terminal-Bench only | Independent | $4.38 | 5.21 |
| #48 | Mistral Small 4Mistral· non-default tier | 21.0 | not published | 21.0 | Terminal-Bench only | Independent | $0.263 | 80.00 |
| #49 | GPT-OSS 120BOpenAI· non-default tier | 13.9 | not published | 13.9 | Terminal-Bench only | Independent | $0.077 | 180.99 |
| #50 | Mistral Large 3Mistral | 12.0 | not published | 12.0 | Terminal-Bench only | Independent | $0.750 | 16.00 |
| #51 | GPT-4.1 MiniOpenAI | 10.1 | not published | 10.1 | Terminal-Bench only | Independent | $0.700 | 14.43 |
| #52 | Ministral 3 14BMistral | 9.7 | not published | 9.7 | Terminal-Bench only | Independent | $0.200 | 48.50 |
| #53 | Llama 4 MaverickMeta | 7.9 | not published | 7.9 | Terminal-Bench only | Independent | $0.415 | 19.04 |
| #54 | GPT-4o MiniOpenAI | 5.6 | not published | 5.6 | Terminal-Bench only | Independent | $0.263 | 21.33 |
| #55 | Ministral 3 8BMistral | 4.1 | not published | 4.1 | Terminal-Bench only | Independent | $0.150 | 27.33 |
| #56 | Llama 4 ScoutMeta | 3.7 | not published | 3.7 | Terminal-Bench only | Independent | $0.135 | 27.41 |
| #57 | GPT-4.1 NanoOpenAI | 3.7 | not published | 3.7 | Terminal-Bench only | Independent | $0.175 | 21.14 |