Skip to main content
TokenCost logoTokenCost

The Best LLMs on SWE-bench Verified

Which model solves the most SWE-bench Verified issues?

Only 7 of the 144 buyable models publish a SWE-bench Verified number at all. The other 137 are named in full below, with whatever they publish instead.

Method: Published SWE-bench Verified pass@1, descending, exactly as reported. Every figure here comes from the provider's own announcement, because no independent lab publishes this suite across the catalogue. Pricing as of July 2026. 1 row below carries a score measured at a non-default reasoning tier and is marked as such; every affected model is named on the hub. Read the full method.

RankModelSWE-bench VerifiedTerminal-BenchQualityReported byBlended $/1M
#1Claude Opus 4.8Anthropic88.684.687.0Unstated / mixed$10.00
#2Claude Sonnet 5Anthropic· non-default tier85.280.583.3Unstated / mixed$4.00
#3Nex-N2-ProNex AGI80.867.875.6Unstated / mixed$1.00
#4MiMo-V2-ProXiaomi78.0not published78.0Unstated / mixed$1.50
#5Mistral Medium 3.5Mistral77.650.666.8Unstated / mixed$3.00
#6Nemotron 3 Ultra 550BNVIDIA71.9not published71.9Unstated / mixed$1.00
#7Qwen3 Coder NextAlibaba70.638.257.6Unstated / mixed$0.282

The 137 models that publish no SWE-bench Verified score

This list is short because the suite is rarely reported, not because the catalogue is small. Every buyable model without a published SWE-bench Verified number is named below, with whatever it does publish instead. An absent score is a fact about what a provider chose to release, and none of these are estimated to fill the table out.

OpenAI

33
  • GPT-4.1GPQA 66.6 · tau2 47.1 · $3.50/1M
  • GPT-4.1 MiniTerminal-Bench 2.1 10.1 · GPQA 66.4 · tau2 52.9 · $0.700/1M
  • GPT-4.1 NanoTerminal-Bench 2.1 3.7 · GPQA 51.2 · tau2 17.3 · $0.175/1M
  • GPT-4oGPQA 54.3 · tau2 25.1 · $4.38/1M
  • GPT-4o MiniTerminal-Bench 2.1 5.6 · GPQA 42.6 · $0.263/1M
  • GPT-5GPQA 84.2 · tau2 86.5 · $3.44/1M
  • GPT-5 MediumGPQA 84.2 · tau2 86.5 · $3.44/1M
  • GPT-5 MiniGPQA 80.3 · tau2 71.1 · $0.688/1M
  • GPT-5 NanoGPQA 67.0 · tau2 30.4 · $0.138/1M
  • GPT-5.1Terminal-Bench 2.1 52.4 · GPQA 87.3 · tau2 81.9 · $3.44/1M
  • GPT-5.2GPQA 86.4 · tau2 74.3 · $4.81/1M
  • GPT-5.3 CodexGPQA 91.5 · tau2 86.0 · $4.81/1M
  • GPT-5.4GPQA 87.1 · tau2 74.6 · $5.63/1M
  • GPT-5.4 MiniGPQA 82.3 · tau2 36.5 · $1.69/1M
  • GPT-5.4 NanoGPQA 76.1 · tau2 52.6 · $0.463/1M
  • GPT-5.4 Pronothing published · $67.50/1M
  • GPT-5.5Terminal-Bench 2.1 80.5 · GPQA 92.6 · tau2 91.8 · $11.25/1M
  • GPT-5.5 Pronothing published · $67.50/1M
  • GPT-5.6 LunaTerminal-Bench 2.1 53.2 · GPQA 85.9 · $2.25/1M
  • GPT-5.6 SolTerminal-Bench 2.1 86.1 · GPQA 92.6 · tau2 81.0 · $11.25/1M
  • GPT-5.6 TerraTerminal-Bench 2.1 72.3 · GPQA 87.2 · tau2 72.8 · $5.63/1M
  • GPT-OSS 120BTerminal-Bench 2.1 13.9 · GPQA 67.2 · tau2 45.0 · $0.077/1M
  • GPT-OSS 120B (Bedrock)nothing published · $0.263/1M
  • GPT-OSS 20BGPQA 61.1 · tau2 50.3 · $0.131/1M
  • GPT-OSS 20B (Bedrock)nothing published · $0.128/1M
  • o1GPQA 74.7 · tau2 62.6 · $26.25/1M
  • o1 ProAA index 18.9 · $262.50/1M
  • o3GPQA 82.7 · tau2 80.7 · $3.50/1M
  • o3 Deep Researchnothing published · $17.50/1M
  • o3 MiniGPQA 74.8 · tau2 28.7 · $1.93/1M
  • o3-proGPQA 84.5 · $35.00/1M
  • o4 MiniGPQA 78.4 · tau2 55.6 · $1.93/1M
  • o4 Mini Deep Researchnothing published · $3.50/1M

Google

17

Mistral

14

xAI

13

Alibaba

11

Anthropic

10

Moonshot

6
  • Kimi K2 ThinkingGPQA 83.8 · tau2 93.0 · $1.07/1M
  • Kimi K2 Thinking Turbonothing published · $2.86/1M
  • Kimi K2.5Terminal-Bench 2.1 45.7 · GPQA 87.9 · tau2 95.9 · $1.20/1M
  • Kimi K2.6Terminal-Bench 2.1 65.9 · GPQA 91.1 · tau2 95.9 · $1.71/1M
  • Kimi K2.7 CodeTerminal-Bench 2.1 67.4 · GPQA 89.6 · tau2 90.1 · $1.71/1M
  • Kimi K3Terminal-Bench 2.1 85.0 · GPQA 93.5 · $6.00/1M

Zhipu

6
  • GLM-4.7Terminal-Bench 2.1 45.3 · GPQA 85.9 · tau2 95.9 · $1.00/1M
  • GLM-4.7-flashGPQA 58.1 · tau2 98.8 · $0.152/1M
  • GLM-5GPQA 82.0 · tau2 98.2 · $1.55/1M
  • GLM-5 TurboGPQA 84.7 · tau2 98.5 · $1.90/1M
  • GLM-5.1Terminal-Bench 2.1 61.8 · GPQA 86.8 · tau2 97.7 · $2.15/1M
  • GLM-5.2Terminal-Bench 2.1 77.9 · GPQA 89.5 · tau2 99.1 · $2.15/1M

DeepSeek

5

Meta

4

MiniMax

3
  • MiniMax M2.5GPQA 84.8 · tau2 95.3 · $0.525/1M
  • MiniMax M2.7Terminal-Bench 2.1 55.4 · GPQA 87.4 · tau2 84.8 · $0.525/1M
  • MiniMax M3Terminal-Bench 2.1 65.2 · GPQA 92.9 · tau2 88.9 · $1.05/1M

Amazon

2

ByteDance

2

Cohere

2
  • Command ATerminal-Bench 2.1 22.8 · GPQA 76.1 · tau2 80.7 · $4.38/1M
  • Command A+Terminal-Bench 2.1 22.8 · GPQA 76.1 · tau2 80.7 · $4.38/1M

Kwaipilot

2

Inception Labs

1
  • Mercury 2Terminal-Bench 2.1 27.3 · GPQA 77.0 · tau2 70.8 · $0.375/1M

Meituan

1
  • LongCat-2.0Terminal-Bench 2.1 50.2 · GPQA 78.0 · $1.30/1M

Microsoft

1

NVIDIA

1

Perplexity

1

StepFun

1
  • Step 3.7 FlashTerminal-Bench 2.1 39.3 · GPQA 80.9 · tau2 98.5 · $0.438/1M

Xiaomi

1
  • MiMo-V2.5-ProTerminal-Bench 2.1 65.2 · GPQA 86.6 · tau2 94.2 · $0.544/1M

Frequently asked

Nearby questions, answered elsewhere

All rankings