The Best LLMs on Terminal-Bench Hard
Which model holds up on the hard Terminal-Bench subset?
86 of the 144 buyable models publish a Terminal-Bench Hard pass rate — more than publish Terminal-Bench 2.1 (55). 43 of them publish no 2.1 figure at all, so this is the only pass-rate list on the site that ranks them.
Method: Published Terminal-Bench Hard pass rate, descending, exactly as reported by Artificial Analysis. Hard and 2.1 are never averaged together: they are different task sets scored on different populations of models, and a mean of the two would be an invented number. Pricing as of July 2026. 31 rows below carry a score measured at a non-default reasoning tier and are marked as such; every affected model is named on the hub. Read the full method.
| Rank | Model | Terminal-Bench Hard | Terminal-Bench | AA coding | Reported by | Blended $/1M | Context |
|---|---|---|---|---|---|---|---|
| #1 | GPT-5.6 SolOpenAI | 62.9 | 86.1 | 76.3 | Independent | $11.25 | 1.05M |
| #2 | Claude Fable 5Anthropic | 62.9 | 84.6 | 76.5 | Independent | $20.00 | 1M |
| #3 | Claude Opus 4.8Anthropic | 58.3 | 84.6 | 74.3 | Unstated / mixed | $10.00 | 1M |
| #4 | GPT-5.5OpenAI | 57.6 | 80.5 | 71.5 | Independent | $11.25 | 1.05M |
| #5 | GPT-5.3 CodexOpenAI | 53.0 | not published | not published | Unstated / mixed | $4.81 | 400K |
| #6 | Claude Sonnet 4.6Anthropic· non-default tier | 53.0 | 71.2 | 63.0 | Independent | $6.00 | 1M |
| #7 | Claude Opus 4.7Anthropic· non-default tier | 51.5 | 83.1 | 73.6 | Independent | $10.00 | 1M |
| #8 | GLM-5.2Zhipu· non-default tier | 50.8 | 77.9 | 68.8 | Independent | $2.15 | 1M |
| #9 | Qwen3.7 MaxAlibaba | 50.8 | 74.5 | 66.0 | Independent | $3.75 | 1M |
| #10 | Qwen3.7 PlusAlibaba | 47.0 | 61.0 | 55.9 | Independent | $0.700 | 1M |
| #11 | Claude Opus 4.5Anthropic· non-default tier | 47.0 | not published | not published | Unstated / mixed | $10.00 | 200K |
| #12 | Claude Opus 4.6Anthropic· non-default tier | 46.2 | not published | not published | Unstated / mixed | $10.00 | 1M |
| #13 | GPT-5.1OpenAI· non-default tier | 45.5 | 52.4 | 49.4 | Independent | $3.44 | 400K |
| #14 | Kimi K2.7 CodeMoonshot | 44.7 | 67.4 | 60.8 | Independent | $1.71 | 262K |
| #15 | Qwen3.6-PlusAlibaba | 43.9 | 61.4 | 54.5 | Independent | $0.620 | 1M |
| #16 | Kimi K2.6Moonshot | 43.9 | 65.9 | 61.8 | Independent | $1.71 | 262K |
| #17 | Qwen3.6-Max-PreviewAlibaba | 43.9 | not published | not published | Unstated / mixed | $2.92 | 262K |
| #18 | MiMo-V2.5-ProXiaomi | 43.2 | 65.2 | 60.2 | Independent | $0.544 | 1.05M |
| #19 | GLM-5Zhipu· non-default tier | 43.2 | not published | not published | Unstated / mixed | $1.55 | 128K |
| #20 | GLM-5.1Zhipu· non-default tier | 43.2 | 61.8 | 55.8 | Independent | $2.15 | 200K |
| #21 | GPT-5.2OpenAI | 43.2 | not published | not published | Unstated / mixed | $4.81 | 400K |
| #22 | GPT-5.4OpenAI· non-default tier | 43.2 | not published | not published | Unstated / mixed | $5.63 | 1.05M |
| #23 | MiniMax M3MiniMax | 42.4 | 65.2 | 58.6 | Independent | $1.05 | 1M |
| #24 | DeepSeek V4-ProDeepSeek· non-default tier | 41.7 | 64.8 | 58.7 | Independent | $0.544 | 1M |
| #25 | MiMo-V2-ProXiaomi | 40.9 | not published | not published | Unstated / mixed | $1.50 | 1.05M |
| #26 | MiniMax M2.7MiniMax | 39.4 | 55.4 | 52.6 | Independent | $0.525 | 205K |
| #27 | Gemini 3.5 FlashGoogle | 39.4 | 76.2 | not published | Unstated / mixed | $3.38 | 1.05M |
| #28 | DeepSeek V4-FlashDeepSeek· non-default tier | 38.6 | 56.9 | 52.0 | Independent | $0.175 | 1M |
| #29 | Gemini 3 Flash ReasoningGoogle· non-default tier | 38.6 | not published | not published | Unstated / mixed | $1.13 | 1.05M |
| #30 | Grok 4.20xAI· non-default tier | 37.9 | not published | not published | Unstated / mixed | $3.00 | 2M |
| #31 | GPT-5OpenAI | 37.9 | not published | not published | Unstated / mixed | $3.44 | 400K |
| #32 | GPT-5 MediumOpenAI | 37.9 | not published | not published | Unstated / mixed | $3.44 | 400K |
| #33 | Grok 4xAI | 37.9 | not published | not published | Unstated / mixed | $6.00 | 2M |
| #34 | o3OpenAI | 37.1 | not published | not published | Unstated / mixed | $3.50 | 200K |
| #35 | Gemma 4 31BGoogle· non-default tier | 36.4 | 43.4 | 43.4 | Independent | $0.180 | 262K |
| #36 | Step 3.7 FlashStepFun | 35.6 | 39.3 | 39.6 | Independent | $0.438 | 262K |
| #37 | MiniMax M2.5MiniMax | 34.8 | not published | not published | Unstated / mixed | $0.525 | 128K |
| #38 | Kimi K2.5Moonshot· non-default tier | 34.8 | 45.7 | 46.8 | Independent | $1.20 | 262K |
| #39 | GPT-5.4 MiniOpenAI | 34.1 | not published | not published | Unstated / mixed | $1.69 | 400K |
| #40 | GPT-5.4 NanoOpenAI | 33.3 | not published | not published | Unstated / mixed | $0.463 | 400K |
| #41 | GLM-5 TurboZhipu | 33.3 | not published | not published | Unstated / mixed | $1.90 | 200K |
| #42 | Mistral Medium 3.5Mistral | 33.3 | 50.6 | 46.9 | Unstated / mixed | $3.00 | 256K |
| #43 | Qwen 3.5 27BAlibaba· non-default tier | 32.6 | not published | not published | Unstated / mixed | $0.825 | 128K |
| #44 | GLM-4.7Zhipu· non-default tier | 31.8 | 45.3 | 45.3 | Independent | $1.00 | 200K |
| #45 | Kimi K2 ThinkingMoonshot | 31.1 | not published | not published | Unstated / mixed | $1.07 | 262K |
| #46 | Grok 4.3xAI | 30.3 | not published | not published | Unstated / mixed | $1.56 | 1M |
| #47 | GPT-5 MiniOpenAI | 28.8 | not published | not published | Unstated / mixed | $0.688 | 400K |
| #48 | Mercury 2Inception Labs | 26.5 | 27.3 | 31.1 | Independent | $0.375 | 128K |
| #49 | Gemini 2.5 ProGoogle | 26.5 | 28.5 | 33.3 | Independent | $3.44 | 1.05M |
| #50 | Command ACohere | 25.0 | 22.8 | 27.8 | Independent | $4.38 | 128K |
| #51 | Command A+Cohere | 25.0 | 22.8 | 27.8 | Independent | $4.38 | 128K |
| #52 | Qwen3.5-9BAlibaba· non-default tier | 24.2 | 29.2 | 28.7 | Independent | $0.075 | 262K |
| #53 | Grok 4.1 FastxAI· non-default tier | 24.2 | not published | not published | Unstated / mixed | $0.275 | 2M |
| #54 | Grok 4.1 Fast ReasoningxAI· non-default tier | 24.2 | not published | not published | Unstated / mixed | $0.275 | 2M |
| #55 | Gemini 3.1 Flash-LiteGoogle | 24.2 | 31.1 | 34.7 | Independent | $0.563 | 1.05M |
| #56 | GLM-4.7-flashZhipu· non-default tier | 22.0 | not published | not published | Unstated / mixed | $0.152 | 200K |
| #57 | Qwen3.5-Omni PlusAlibaba | 21.2 | not published | not published | Unstated / mixed | $1.50 | 262K |
| #58 | Claude 3.7 SonnetAnthropic· non-default tier | 21.2 | not published | 36.4 | Unstated / mixed | $6.00 | 200K |
| #59 | Gemma 4 12BGoogle· non-default tier | 18.2 | 27.3 | 31.0 | Independent | $0.00 | 262K |
| #60 | Qwen3 Coder NextAlibaba | 18.2 | 38.2 | 36.2 | Unstated / mixed | $0.282 | 262K |
| #61 | GPT-5 NanoOpenAI | 17.4 | not published | not published | Unstated / mixed | $0.138 | 128K |
| #62 | Mistral Small 4Mistral· non-default tier | 17.4 | 21.0 | 26.6 | Independent | $0.263 | 256K |
| #63 | Nova 2.0 LiteAmazon | 17.4 | not published | not published | Unstated / mixed | $0.850 | 1M |
| #64 | Mistral Large 3Mistral | 15.9 | 12.0 | 20.1 | Independent | $0.750 | 262K |
| #65 | DeepSeek R1DeepSeek | 15.9 | not published | not published | Unstated / mixed | $0.960 | 128K |
| #66 | o4 MiniOpenAI | 15.2 | not published | not published | Unstated / mixed | $1.93 | 200K |
| #67 | Gemma 4 26B A4BGoogle· non-default tier | 13.6 | 39.0 | 39.3 | Independent | $0.120 | 262K |
| #68 | Gemini 2.5 FlashGoogle· non-default tier | 13.6 | not published | not published | Unstated / mixed | $0.850 | 1.05M |
| #69 | GPT-4.1OpenAI | 13.6 | not published | not published | Unstated / mixed | $3.50 | 1.05M |
| #70 | o1OpenAI | 12.9 | not published | 39.7 | Unstated / mixed | $26.25 | 200K |
| #71 | Grok 3xAI | 11.4 | not published | not published | Unstated / mixed | $6.00 | 131K |
| #72 | Magistral MediumMistral | 9.1 | not published | not published | Unstated / mixed | $2.75 | 40K |
| #73 | Qwen3.5-Omni FlashAlibaba | 8.3 | not published | not published | Unstated / mixed | $0.850 | 262K |
| #74 | GPT-4oOpenAI· non-default tier | 8.3 | not published | not published | Unstated / mixed | $4.38 | 128K |
| #75 | GPT-4.1 MiniOpenAI | 7.6 | 10.1 | 20.2 | Independent | $0.700 | 1.05M |
| #76 | Llama 4 MaverickMeta | 6.8 | 7.9 | 16.3 | Independent | $0.415 | 1.05M |
| #77 | o3 MiniOpenAI | 6.8 | not published | not published | Unstated / mixed | $1.93 | 200K |
| #78 | Devstral SmallMistral· non-default tier | 6.1 | not published | not published | Unstated / mixed | $0.150 | 256K |
| #79 | GPT-OSS 120BOpenAI· non-default tier | 5.3 | 13.9 | 21.2 | Independent | $0.077 | 131K |
| #80 | GPT-OSS 20BOpenAI· non-default tier | 4.5 | not published | not published | Unstated / mixed | $0.131 | 131K |
| #81 | Ministral 3 8BMistral | 4.5 | 4.1 | 9.7 | Independent | $0.150 | 262K |
| #82 | Gemini 2.5 Flash-LiteGoogle· non-default tier | 4.5 | not published | not published | Unstated / mixed | $0.175 | 1.05M |
| #83 | Ministral 3 14BMistral | 4.5 | 9.7 | 14.4 | Independent | $0.200 | 262K |
| #84 | Magistral SmallMistral | 4.5 | not published | not published | Unstated / mixed | $0.750 | 40K |
| #85 | GPT-4.1 NanoOpenAI | 3.8 | 3.7 | 11.1 | Independent | $0.175 | 1.05M |
| #86 | Llama 4 ScoutMeta | 1.5 | 3.7 | 8.2 | Independent | $0.135 | 1.05M |
Frequently asked
Nearby questions, answered elsewhere
All rankings
The hub and the methodThe Best LLMs by Benchmark ScoreThe Best Value LLMs per DollarThe Best LLMs by Context WindowThe Cheapest LLMs With a 1M-Token Context WindowThe Best LLMs on Terminal-Bench 2.1The Best LLMs on SWE-bench VerifiedThe Best LLMs on GPQA DiamondThe Best LLMs for Tool Use on tau2-bench