The Best LLMs on Terminal-Bench 2.1
Which model solves the most Terminal-Bench 2.1 tasks?
55 of the 144 buyable models publish a Terminal-Bench 2.1 pass rate. 86 also publish Terminal-Bench Hard, shown alongside.
Method: Published Terminal-Bench 2.1 pass rate, descending, exactly as reported. Nothing is weighted, blended or rescaled here: this is one suite, reported as published, with price alongside it rather than folded into it. Pricing as of July 2026. 16 rows below carry a score measured at a non-default reasoning tier and are marked as such; every affected model is named on the hub. Read the full method.
| Rank | Model | Terminal-Bench 2.1 | TB Hard | AA coding | Reported by | Blended $/1M | Context |
|---|---|---|---|---|---|---|---|
| #1 | GPT-5.6 SolOpenAI | 86.1 | 62.9 | 76.3 | Independent | $11.25 | 1.05M |
| #2 | Kimi K3Moonshot | 85.0 | not published | 76.2 | Independent | $6.00 | 1.05M |
| #3 | Claude Opus 4.8Anthropic | 84.6 | 58.3 | 74.3 | Unstated / mixed | $10.00 | 1M |
| #4 | Claude Fable 5Anthropic | 84.6 | 62.9 | 76.5 | Independent | $20.00 | 1M |
| #5 | Claude Opus 4.7Anthropic· non-default tier | 83.1 | 51.5 | 73.6 | Independent | $10.00 | 1M |
| #6 | Grok 4.5xAI | 81.6 | not published | 72.4 | Independent | $3.00 | 500K |
| #7 | Claude Sonnet 5Anthropic· non-default tier | 80.5 | not published | 71.5 | Unstated / mixed | $4.00 | 1M |
| #8 | GPT-5.5OpenAI | 80.5 | 57.6 | 71.5 | Independent | $11.25 | 1.05M |
| #9 | Muse Spark 1.1Meta | 77.9 | not published | 71.3 | Independent | $2.00 | 1M |
| #10 | GLM-5.2Zhipu· non-default tier | 77.9 | 50.8 | 68.8 | Independent | $2.15 | 1M |
| #11 | Gemini 3.6 FlashGoogle | 77.5 | not published | 69.2 | Independent | $3.00 | 1.05M |
| #12 | Gemini 3.5 FlashGoogle | 76.2 | 39.4 | not published | Unstated / mixed | $3.38 | 1.05M |
| #13 | Qwen3.7 MaxAlibaba | 74.5 | 50.8 | 66.0 | Independent | $3.75 | 1M |
| #14 | GPT-5.6 TerraOpenAI | 72.3 | not published | 64.7 | Independent | $5.63 | 1.05M |
| #15 | Claude Sonnet 4.6Anthropic· non-default tier | 71.2 | 53.0 | 63.0 | Independent | $6.00 | 1M |
| #16 | Nex-N2-ProNex AGI | 67.8 | not published | 59.1 | Unstated / mixed | $1.00 | 262K |
| #17 | Kimi K2.7 CodeMoonshot | 67.4 | 44.7 | 60.8 | Independent | $1.71 | 262K |
| #18 | Kimi K2.6Moonshot | 65.9 | 43.9 | 61.8 | Independent | $1.71 | 262K |
| #19 | MiMo-V2.5-ProXiaomi | 65.2 | 43.2 | 60.2 | Independent | $0.544 | 1.05M |
| #20 | MiniMax M3MiniMax | 65.2 | 42.4 | 58.6 | Independent | $1.05 | 1M |
| #21 | DeepSeek V4-ProDeepSeek· non-default tier | 64.8 | 41.7 | 58.7 | Independent | $0.544 | 1M |
| #22 | GLM-5.1Zhipu· non-default tier | 61.8 | 43.2 | 55.8 | Independent | $2.15 | 200K |
| #23 | Qwen3.6-PlusAlibaba | 61.4 | 43.9 | 54.5 | Independent | $0.620 | 1M |
| #24 | Qwen3.7 PlusAlibaba | 61.0 | 47.0 | 55.9 | Independent | $0.700 | 1M |
| #25 | KAT-Coder-Pro v2.5Kwaipilot | 60.7 | not published | not published | Vendor-reported | $1.29 | 262K |
| #26 | DeepSeek V4-FlashDeepSeek· non-default tier | 56.9 | 38.6 | 52.0 | Independent | $0.175 | 1M |
| #27 | MiniMax M2.7MiniMax | 55.4 | 39.4 | 52.6 | Independent | $0.525 | 205K |
| #28 | Gemini 3.5 Flash-LiteGoogle | 53.6 | not published | 49.3 | Independent | $0.850 | 1.05M |
| #29 | GPT-5.6 LunaOpenAI | 53.2 | not published | 50.7 | Independent | $2.25 | 1.05M |
| #30 | GPT-5.1OpenAI· non-default tier | 52.4 | 45.5 | 49.4 | Independent | $3.44 | 400K |
| #31 | Mistral Medium 3.5Mistral | 50.6 | 33.3 | 46.9 | Unstated / mixed | $3.00 | 256K |
| #32 | LongCat-2.0Meituan | 50.2 | not published | 45.3 | Independent | $1.30 | 1M |
| #33 | Kimi K2.5Moonshot· non-default tier | 45.7 | 34.8 | 46.8 | Independent | $1.20 | 262K |
| #34 | GLM-4.7Zhipu· non-default tier | 45.3 | 31.8 | 45.3 | Independent | $1.00 | 200K |
| #35 | Gemma 4 31BGoogle· non-default tier | 43.4 | 36.4 | 43.4 | Independent | $0.180 | 262K |
| #36 | Step 3.7 FlashStepFun | 39.3 | 35.6 | 39.6 | Independent | $0.438 | 262K |
| #37 | Gemma 4 26B A4BGoogle· non-default tier | 39.0 | 13.6 | 39.3 | Independent | $0.120 | 262K |
| #38 | Qwen3 Coder NextAlibaba | 38.2 | 18.2 | 36.2 | Unstated / mixed | $0.282 | 262K |
| #39 | Gemini 3.1 Flash-LiteGoogle | 31.1 | 24.2 | 34.7 | Independent | $0.563 | 1.05M |
| #40 | Qwen3.5-9BAlibaba· non-default tier | 29.2 | 24.2 | 28.7 | Independent | $0.075 | 262K |
| #41 | Gemini 2.5 ProGoogle | 28.5 | 26.5 | 33.3 | Independent | $3.44 | 1.05M |
| #42 | Gemma 4 12BGoogle· non-default tier | 27.3 | 18.2 | 31.0 | Independent | $0.00 | 262K |
| #43 | Mercury 2Inception Labs | 27.3 | 26.5 | 31.1 | Independent | $0.375 | 128K |
| #44 | Command ACohere | 22.8 | 25.0 | 27.8 | Independent | $4.38 | 128K |
| #45 | Command A+Cohere | 22.8 | 25.0 | 27.8 | Independent | $4.38 | 128K |
| #46 | Mistral Small 4Mistral· non-default tier | 21.0 | 17.4 | 26.6 | Independent | $0.263 | 256K |
| #47 | GPT-OSS 120BOpenAI· non-default tier | 13.9 | 5.3 | 21.2 | Independent | $0.077 | 131K |
| #48 | Mistral Large 3Mistral | 12.0 | 15.9 | 20.1 | Independent | $0.750 | 262K |
| #49 | GPT-4.1 MiniOpenAI | 10.1 | 7.6 | 20.2 | Independent | $0.700 | 1.05M |
| #50 | Ministral 3 14BMistral | 9.7 | 4.5 | 14.4 | Independent | $0.200 | 262K |
| #51 | Llama 4 MaverickMeta | 7.9 | 6.8 | 16.3 | Independent | $0.415 | 1.05M |
| #52 | GPT-4o MiniOpenAI | 5.6 | not published | 11.4 | Independent | $0.263 | 128K |
| #53 | Ministral 3 8BMistral | 4.1 | 4.5 | 9.7 | Independent | $0.150 | 262K |
| #54 | Llama 4 ScoutMeta | 3.7 | 1.5 | 8.2 | Independent | $0.135 | 1.05M |
| #55 | GPT-4.1 NanoOpenAI | 3.7 | 3.8 | 11.1 | Independent | $0.175 | 1.05M |
Frequently asked
Nearby questions, answered elsewhere
All rankings
The hub and the methodThe Best LLMs by Benchmark ScoreThe Best Value LLMs per DollarThe Best LLMs by Context WindowThe Cheapest LLMs With a 1M-Token Context WindowThe Best LLMs on Terminal-Bench HardThe Best LLMs on SWE-bench VerifiedThe Best LLMs on GPQA DiamondThe Best LLMs for Tool Use on tau2-bench