Skip to main content
TokenCost logoTokenCost

The Best LLMs by Benchmark Score

Which model scores highest on the published coding benchmarks?

57 of the 144 buyable models carry a published pass-rate score. The other 87 are absent from this list rather than ranked at the bottom on a zero they never earned.

Method: Weighted mean of the published pass-rate suites the model has (SWE-bench Verified 0.6, Terminal-Bench 2.1 0.4), renormalised over the suites present. An exact tie breaks first to the model with more suites published, since a score corroborated by two suites is a stronger claim than the same score from one; then to the lower blended price, then model id. Pricing as of July 2026. 16 rows below carry a score measured at a non-default reasoning tier and are marked as such; every affected model is named on the hub. Read the full method.

RankModelQualitySWE-bench V.Terminal-BenchEvidenceReported byBlended $/1MPer $
#1Claude Opus 4.8Anthropic87.088.684.6Both suitesUnstated / mixed$10.008.70
#2GPT-5.6 SolOpenAI86.1not published86.1Terminal-Bench onlyIndependent$11.257.65
#3Kimi K3Moonshot85.0not published85.0Terminal-Bench onlyIndependent$6.0014.17
#4Claude Fable 5Anthropic84.6not published84.6Terminal-Bench onlyIndependent$20.004.23
#5Claude Sonnet 5Anthropic· non-default tier83.385.280.5Both suitesUnstated / mixed$4.0020.83
#6Claude Opus 4.7Anthropic· non-default tier83.1not published83.1Terminal-Bench onlyIndependent$10.008.31
#7Grok 4.5xAI81.6not published81.6Terminal-Bench onlyIndependent$3.0027.20
#8GPT-5.5OpenAI80.5not published80.5Terminal-Bench onlyIndependent$11.257.16
#9MiMo-V2-ProXiaomi78.078.0not publishedSWE-bench onlyUnstated / mixed$1.5052.00
#10Muse Spark 1.1Meta77.9not published77.9Terminal-Bench onlyIndependent$2.0038.95
#11GLM-5.2Zhipu· non-default tier77.9not published77.9Terminal-Bench onlyIndependent$2.1536.23
#12Gemini 3.6 FlashGoogle77.5not published77.5Terminal-Bench onlyIndependent$3.0025.83
#13Gemini 3.5 FlashGoogle76.2not published76.2Terminal-Bench onlyUnstated / mixed$3.3822.58
#14Nex-N2-ProNex AGI75.680.867.8Both suitesUnstated / mixed$1.0075.60
#15Qwen3.7 MaxAlibaba74.5not published74.5Terminal-Bench onlyIndependent$3.7519.87
#16GPT-5.6 TerraOpenAI72.3not published72.3Terminal-Bench onlyIndependent$5.6312.85
#17Nemotron 3 Ultra 550BNVIDIA71.971.9not publishedSWE-bench onlyUnstated / mixed$1.0071.90
#18Claude Sonnet 4.6Anthropic· non-default tier71.2not published71.2Terminal-Bench onlyIndependent$6.0011.87
#19Kimi K2.7 CodeMoonshot67.4not published67.4Terminal-Bench onlyIndependent$1.7139.36
#20Mistral Medium 3.5Mistral66.877.650.6Both suitesUnstated / mixed$3.0022.27
#21Kimi K2.6Moonshot65.9not published65.9Terminal-Bench onlyIndependent$1.7138.48
#22MiMo-V2.5-ProXiaomi65.2not published65.2Terminal-Bench onlyIndependent$0.544119.90
#23MiniMax M3MiniMax65.2not published65.2Terminal-Bench onlyIndependent$1.0562.10
#24DeepSeek V4-ProDeepSeek· non-default tier64.8not published64.8Terminal-Bench onlyIndependent$0.544119.16
#25GLM-5.1Zhipu· non-default tier61.8not published61.8Terminal-Bench onlyIndependent$2.1528.74
#26Qwen3.6-PlusAlibaba61.4not published61.4Terminal-Bench onlyIndependent$0.62099.06
#27Qwen3.7 PlusAlibaba61.0not published61.0Terminal-Bench onlyIndependent$0.70087.14
#28KAT-Coder-Pro v2.5Kwaipilot60.7not published60.7Terminal-Bench onlyVendor-reported$1.2946.87
#29Qwen3 Coder NextAlibaba57.670.638.2Both suitesUnstated / mixed$0.282203.89
#30DeepSeek V4-FlashDeepSeek· non-default tier56.9not published56.9Terminal-Bench onlyIndependent$0.175325.14
#31MiniMax M2.7MiniMax55.4not published55.4Terminal-Bench onlyIndependent$0.525105.52
#32Gemini 3.5 Flash-LiteGoogle53.6not published53.6Terminal-Bench onlyIndependent$0.85063.06
#33GPT-5.6 LunaOpenAI53.2not published53.2Terminal-Bench onlyIndependent$2.2523.64
#34GPT-5.1OpenAI· non-default tier52.4not published52.4Terminal-Bench onlyIndependent$3.4415.24
#35LongCat-2.0Meituan50.2not published50.2Terminal-Bench onlyIndependent$1.3038.62
#36Kimi K2.5Moonshot· non-default tier45.7not published45.7Terminal-Bench onlyIndependent$1.2038.08
#37GLM-4.7Zhipu· non-default tier45.3not published45.3Terminal-Bench onlyIndependent$1.0045.30
#38Gemma 4 31BGoogle· non-default tier43.4not published43.4Terminal-Bench onlyIndependent$0.180241.11
#39Step 3.7 FlashStepFun39.3not published39.3Terminal-Bench onlyIndependent$0.43889.83
#40Gemma 4 26B A4BGoogle· non-default tier39.0not published39.0Terminal-Bench onlyIndependent$0.120325.00
#41Gemini 3.1 Flash-LiteGoogle31.1not published31.1Terminal-Bench onlyIndependent$0.56355.29
#42Qwen3.5-9BAlibaba· non-default tier29.2not published29.2Terminal-Bench onlyIndependent$0.075389.33
#43Gemini 2.5 ProGoogle28.5not published28.5Terminal-Bench onlyIndependent$3.448.29
#44Gemma 4 12BGoogle· non-default tier27.3not published27.3Terminal-Bench onlyIndependent$0.00
#45Mercury 2Inception Labs27.3not published27.3Terminal-Bench onlyIndependent$0.37572.80
#46Command ACohere22.8not published22.8Terminal-Bench onlyIndependent$4.385.21
#47Command A+Cohere22.8not published22.8Terminal-Bench onlyIndependent$4.385.21
#48Mistral Small 4Mistral· non-default tier21.0not published21.0Terminal-Bench onlyIndependent$0.26380.00
#49GPT-OSS 120BOpenAI· non-default tier13.9not published13.9Terminal-Bench onlyIndependent$0.077180.99
#50Mistral Large 3Mistral12.0not published12.0Terminal-Bench onlyIndependent$0.75016.00
#51GPT-4.1 MiniOpenAI10.1not published10.1Terminal-Bench onlyIndependent$0.70014.43
#52Ministral 3 14BMistral9.7not published9.7Terminal-Bench onlyIndependent$0.20048.50
#53Llama 4 MaverickMeta7.9not published7.9Terminal-Bench onlyIndependent$0.41519.04
#54GPT-4o MiniOpenAI5.6not published5.6Terminal-Bench onlyIndependent$0.26321.33
#55Ministral 3 8BMistral4.1not published4.1Terminal-Bench onlyIndependent$0.15027.33
#56Llama 4 ScoutMeta3.7not published3.7Terminal-Bench onlyIndependent$0.13527.41
#57GPT-4.1 NanoOpenAI3.7not published3.7Terminal-Bench onlyIndependent$0.17521.14

Frequently asked

Nearby questions, answered elsewhere

All rankings