Best LLM for Coding
A curated comparison of the top LLMs for software development, with API pricing, context windows, and what makes each model stand out for coding tasks.
Prices on this page last verified against each provider’s own pricing page
Order is computed, not hand-picked. Each model is scored on the published pass-rate suites it has (SWE-bench Verified 0.75, Terminal-Bench 2.1 0.25), renormalised over the suites present, then charged 1 quality point for every doubling of its blended price (3:1 input-to-output tokens). Models with no published score are never given one: they are listed after every scored model, ordered by context window, and labelled as unscored. General software work, so SWE-bench Verified (real GitHub issues, patch applies and tests pass) carries 0.75 and Terminal-Bench 2.1 carries 0.25. Price barely presses on the order here: the question this page answers is which model writes the best code, not which is cheapest. See the full ranking method.
Anthropic's flagship tops published coding benchmarks at 96.0 on SWE-bench Verified, with a full 1M context and careful, structured output that makes it the premium pick for refactoring and code review. It replaced Opus 4.8 on July 24, 2026 at the same $5/$25, so there is no reason to pick 4.8 over it on price. A gated Fast mode at $10/$50 buys up to 2.5x output tokens per second.
Moonshot's brand-new flagship, launched July 16, 2026. Ranked #1 on Arena Frontend Code and #4 on the Artificial Analysis Intelligence Index, with the full 1M context billed flat. It reasons verbosely, so real costs can run higher than the rate card suggests.
The default choice for most coding work. Sonnet 5 scores 85.2 on SWE-bench Verified, close behind the Opus tier at 60% lower list price, and its $2/$10 is now the standard rate after Anthropic cancelled the September 1 increase. Full 1M context with no long-context surcharge.
OpenAI's generally available flagship, ranked #1 on the Artificial Analysis Intelligence Index at its April 2026 launch and strong on agentic coding (Terminal-Bench 2.1 83.4). Note the 2x input surcharge above 272K context. The newer GPT-5.6 family is now generally available and cheaper across the board — Terra at $2/$12 and Luna at $0.20/$1.20.
Strong value with a 1M context and up to 384K output tokens, with a hard stop attached. From 12:00 Beijing time on 14 September 2026 DeepSeek routes every deepseek-v4-pro request to V4.1 Flash and bills it at the Flash card ($0.15/$0.60 off-peak), until a V4.1 Pro with no published date. The figure shown here is the off-peak rate that applies until then; weekday peak windows (01:00-04:00 and 06:00-10:00 UTC) cost exactly double.
Google's 1M-context workhorse at $2/1M input, strong at understanding large codebases and generating structured output. For lighter tasks, Gemini 3 Flash offers a cheaper Google option at $0.50/1M input.
OpenAI's coding-focused model at $1.75/1M input, a cheaper path into the GPT-5 line for pure code generation and editing. The 400K context is smaller than the current 1M flagships but plenty for most projects.
How to Choose the Right Coding LLM
For maximum quality: Claude Opus 5 currently leads published coding benchmarks, with GPT-5.5 close behind and stronger on some agentic evals. The preview-only GPT-5.6 Sol posts higher Terminal-Bench scores but remains gated to a small set of partner orgs.
For budget coding: DeepSeek V4-Pro offers remarkable coding ability for the price: a typical 20K-input, 3K-output coding task costs about four cents at its peak rate and half that off-peak, versus roughly $0.07 on Claude Sonnet 5 and $0.18 on Opus 5. Sonnet 5 is the best mid-tier default at a standard $2/$10.
For large codebases: A 1M-token context is now standard at the top. Claude Opus 5, Claude Sonnet 5, GPT-5.5, Gemini 3.1 Pro, Kimi K3, and DeepSeek V4 all offer it, and Kimi K3 bills the full window flat with no length tiering.
Frequently Asked Questions
Common questions about choosing an LLM for coding