Skip to main content
TokenCost logoTokenCost

The Best LLMs by Benchmark Score

Which model scores highest on the published coding benchmarks?

60 of the 135 buyable models carry a published pass-rate score. The other 75 are absent from this list rather than ranked at the bottom on a zero they never earned.

Method: Weighted mean of the published pass-rate suites the model has (SWE-bench Verified 0.6, Terminal-Bench 2.1 0.4), renormalised over the suites present. An exact tie breaks first to the model with more suites published, since a score corroborated by two suites is a stronger claim than the same score from one; then to the lower blended price, then model id. Pricing as of July 2026. 16 rows below carry a score measured at a non-default reasoning tier and are marked as such; every affected model is named on the hub. Read the full method.

RankModelQualitySWE-bench V.Terminal-BenchEvidenceReported byBlended $/1MPer $
#1Claude Opus 5Anthropic92.096.086.1Both suitesUnstated / mixed$10.009.20
#2Claude Opus 4.8Anthropic87.088.684.6Both suitesUnstated / mixed$10.008.70
#3GPT-5.6 SolOpenAI86.1not published86.1Terminal-Bench onlyIndependent$11.257.65
#4Kimi K3Moonshot85.0not published85.0Terminal-Bench onlyIndependent$6.0014.17
#5Claude Fable 5Anthropic84.6not published84.6Terminal-Bench onlyIndependent$20.004.23
#6Claude Sonnet 5Anthropic· non-default tier83.385.280.5Both suitesUnstated / mixed$4.0020.83
#7Claude Opus 4.7Anthropic· non-default tier83.1not published83.1Terminal-Bench onlyIndependent$10.008.31
#8Grok 4.5xAI81.6not published81.6Terminal-Bench onlyIndependent$3.0027.20
#9GPT-5.5OpenAI80.5not published80.5Terminal-Bench onlyIndependent$11.257.16
#10MiMo-V2-ProXiaomi78.078.0not publishedSWE-bench onlyUnstated / mixed$1.5052.00
#11Muse Spark 1.1Meta77.9not published77.9Terminal-Bench onlyIndependent$2.0038.95
#12GLM-5.2Zhipu· non-default tier77.9not published77.9Terminal-Bench onlyIndependent$2.1536.23
#13Gemini 3.6 FlashGoogle77.5not published77.5Terminal-Bench onlyIndependent$1.5051.67
#14Gemini 3.5 FlashGoogle76.2not published76.2Terminal-Bench onlyUnstated / mixed$3.3822.58
#15Nex-N2-ProNex AGI75.680.867.8Both suitesUnstated / mixed$1.0075.60
#16Qwen3.7 MaxAlibaba74.5not published74.5Terminal-Bench onlyIndependent$3.7519.87
#17GPT-5.6 TerraOpenAI72.3not published72.3Terminal-Bench onlyIndependent$4.5016.07
#18Nemotron 3 Ultra 550BNVIDIA71.971.9not publishedSWE-bench onlyUnstated / mixed$1.0071.90
#19Claude Sonnet 4.6Anthropic· non-default tier71.2not published71.2Terminal-Bench onlyIndependent$6.0011.87
#20Laguna S 2.1Poolside70.2not published70.2Terminal-Bench onlyVendor-reported$0.125561.60
#21Kimi K2.7 CodeMoonshot67.4not published67.4Terminal-Bench onlyIndependent$1.7139.36
#22Mistral Medium 3.5Mistral66.877.650.6Both suitesUnstated / mixed$3.0022.27
#23Muse Glimmer 30BMeta66.376.051.7Both suitesVendor-reported$0.637104.00
#24Kimi K2.6Moonshot65.9not published65.9Terminal-Bench onlyIndependent$1.7138.48
#25MiMo-V2.5-ProXiaomi65.2not published65.2Terminal-Bench onlyIndependent$0.544119.90
#26MiniMax M3MiniMax65.2not published65.2Terminal-Bench onlyIndependent$1.0562.10
#27DeepSeek V4-ProDeepSeek· non-default tier64.8not published64.8Terminal-Bench onlyIndependent$0.544119.16
#28GLM-5.1Zhipu· non-default tier61.8not published61.8Terminal-Bench onlyIndependent$2.1528.74
#29Qwen3.6-PlusAlibaba61.4not published61.4Terminal-Bench onlyIndependent$1.1354.58
#30Qwen3.7 PlusAlibaba61.0not published61.0Terminal-Bench onlyIndependent$0.70087.14
#31KAT-Coder-Pro v2.5Kwaipilot60.7not published60.7Terminal-Bench onlyVendor-reported$1.2946.87
#32Qwen3 Coder NextAlibaba57.670.638.2Both suitesUnstated / mixed$0.282203.89
#33DeepSeek V4-FlashDeepSeek· non-default tier56.9not published56.9Terminal-Bench onlyIndependent$0.175325.14
#34MiniMax M2.7MiniMax55.4not published55.4Terminal-Bench onlyIndependent$0.525105.52
#35Gemini 3.5 Flash-LiteGoogle53.6not published53.6Terminal-Bench onlyIndependent$0.85063.06
#36GPT-5.6 LunaOpenAI53.2not published53.2Terminal-Bench onlyIndependent$0.450118.22
#37GPT-5.1OpenAI· non-default tier52.4not published52.4Terminal-Bench onlyIndependent$3.4415.24
#38LongCat-2.0Meituan50.2not published50.2Terminal-Bench onlyIndependent$1.3038.62
#39Kimi K2.5Moonshot· non-default tier45.7not published45.7Terminal-Bench onlyIndependent$1.2038.08
#40GLM-4.7Zhipu· non-default tier45.3not published45.3Terminal-Bench onlyIndependent$1.0045.30
#41Gemma 4 31BGoogle· non-default tier43.4not published43.4Terminal-Bench onlyIndependent$0.180241.11
#42Step 3.7 FlashStepFun39.3not published39.3Terminal-Bench onlyIndependent$0.43889.83
#43Gemma 4 26B A4BGoogle· non-default tier39.0not published39.0Terminal-Bench onlyIndependent$0.120325.00
#44Gemini 3.1 Flash-LiteGoogle31.1not published31.1Terminal-Bench onlyIndependent$0.56355.29
#45Qwen3.5-9BAlibaba· non-default tier29.2not published29.2Terminal-Bench onlyIndependent$0.113259.56
#46Gemini 2.5 ProGoogle28.5not published28.5Terminal-Bench onlyIndependent$3.448.29
#47Gemma 4 12BGoogle· non-default tier27.3not published27.3Terminal-Bench onlyIndependent$0.00
#48Mercury 2Inception Labs27.3not published27.3Terminal-Bench onlyIndependent$0.37572.80
#49Command ACohere22.8not published22.8Terminal-Bench onlyIndependent$4.385.21
#50Command A+Cohere22.8not published22.8Terminal-Bench onlyIndependent$4.385.21
#51Mistral Small 4Mistral· non-default tier21.0not published21.0Terminal-Bench onlyIndependent$0.26380.00
#52GPT-OSS 120BOpenAI· non-default tier13.9not published13.9Terminal-Bench onlyIndependent$0.077180.99
#53Mistral Large 3Mistral12.0not published12.0Terminal-Bench onlyIndependent$0.75016.00
#54GPT-4.1 MiniOpenAI10.1not published10.1Terminal-Bench onlyIndependent$0.70014.43
#55Ministral 3 14BMistral9.7not published9.7Terminal-Bench onlyIndependent$0.20048.50
#56Llama 4 MaverickMeta7.9not published7.9Terminal-Bench onlyIndependent$0.41519.04
#57GPT-4o MiniOpenAI5.6not published5.6Terminal-Bench onlyIndependent$0.26321.33
#58Ministral 3 8BMistral4.1not published4.1Terminal-Bench onlyIndependent$0.15027.33
#59Llama 4 ScoutMeta3.7not published3.7Terminal-Bench onlyIndependent$0.13527.41
#60GPT-4.1 NanoOpenAI3.7not published3.7Terminal-Bench onlyIndependent$0.17521.14

Frequently asked

Nearby questions, answered elsewhere

All rankings