Skip to main content
TokenCost logoTokenCost

Best LLM for Kilo Code

Kilo Code takes no markup on your key, so the model you pick is the bill you pay. This list is ordered accordingly.

Kilo Code is an open-source agentic coding assistant that runs in VS Code, in JetBrains IDEs and as a terminal CLI, all reading the same kilo.json config. It is bring-your-own-key with zero markup: you connect OpenRouter or a provider directly, and every token bills to you at the provider's rate. Its built-in agents split the work — code and debug get full tool access, plan and ask are read-only — which means most users end up running more than one model and want to know both what to put in the expensive seat and what is good enough for the cheap one.

It is also, by volume, the place where that decision costs the most. OpenRouter's own application leaderboard has Kilo Code at #2 across every snapshot we hold, at 244B tokens in the week to August 3, 2026 and 7.31T over the trailing thirty days, behind only Hermes Agent and ahead of Claude Code. Tools with a fixed backend never appear on that list at all; an app is only there because its users choose their own model. So the reader of this page is not picking a model for one session, they are picking a monthly bill.

That is why this page presses on price harder than any other ranking on the site: every halving of a model's blended price is worth six points of measured pass rate here, against five on our Cline page and one on Copilot, where a subscription absorbs the tokens. The suites are weighted 0.55 SWE-bench Verified to 0.45 Terminal-Bench 2.1, because Kilo Code has to hold up in an editor loop and a shell loop from the same configuration rather than specialising in either. Models below a 200K context window are excluded outright: Kilo Code's middle-out transform will squeeze an over-long prompt to fit, but a model that needs it on every task is not a candidate.

Top Models for Kilo Code in 2026

#1
DeepSeek V4-Flash
DeepSeek
Best for Kilo CodeTerminal-Bench only · 56.9Non-default tier
In: $0.14/1M
Out: $0.28/1M
Ctx: 1.0M

The default this page argues for, and the reason is arithmetic rather than enthusiasm. At $0.14/$0.28 per 1M with a 1M context, V4-Flash is roughly a fourteenth of Sonnet 5's input rate, and its cache hit rate of $0.0028/1M is a 98% discount against the 90% the rest of the industry offers — which matters enormously in an agent loop that re-sends the same file tree on every turn. It scores 56.9 on Terminal-Bench 2.1 and publishes no SWE-bench Verified figure, so it is not the strongest model here; it is the one where being wrong twice still costs less than being right once on a frontier tier. Note the 0731 rebuild shipped under the same model ID and the same rate card.

#2
Claude Sonnet 5
Anthropic
Both suites · 83.1Non-default tier
In: $2/1M
Out: $10/1M
Ctx: 1.0M

The quality pick, and the one to put in the seat where a failed run is expensive. Sonnet 5 scores 85.2 on SWE-bench Verified and 80.5 on Terminal-Bench 2.1 with a full 1M context and no long-context surcharge, and it is still on Anthropic's introductory $2/$10 per 1M — which expires on August 31, 2026, after which it bills $3/$15. Budget for that: a Kilo Code habit sized on the intro rate gets 50% more expensive on September 1 without anyone changing a setting. One quiet cost: Sonnet 5's tokenizer emits roughly 30% more tokens than Sonnet 4.6 for the same text, so the sticker gap understates the real one.

#3
DeepSeek V4-Pro
DeepSeek
Terminal-Bench only · 64.8Non-default tier
In: $0.435/1M
Out: $0.87/1M
Ctx: 1.0M

The step up from V4-Flash without leaving the DeepSeek price band. V4-Pro is a 1.6T-parameter MoE with 49B active and three reasoning modes, scoring 64.8 on Terminal-Bench 2.1 at $0.435/$0.87 — about a fifth of Sonnet 5's input rate for meaningfully harder debugging. Two things to know before you standardise on it: its concurrency limit is 500 against V4-Flash's 2,500, which a parallel Kilo Code workflow can actually hit, and DeepSeek has published but not yet enforced a peak-hour doubling for 09:00-12:00 and 14:00-18:00 Beijing time. The 75% launch discount was made permanent in May, so the base rate is not going to lapse under you.

#4
Qwen3 Coder Next
Alibaba
Both suites · 56.0
In: $0.11/1M
Out: $0.8/1M
Ctx: 262K

The cheapest thing on this page that posts a real SWE-bench Verified number: 70.6, from an Apache 2.0 80B MoE with 3B active, at $0.11/$0.80 per 1M on OpenRouter's cheapest route. For a Kilo Code user whose goal is a monthly bill near zero, this is the model to try first in the code agent. The limits are honest ones — 256K native context against the 1M the Claude and DeepSeek entries carry, non-thinking mode only, and a Terminal-Bench 2.1 score of 38.2 that says it is much better at editing files than at driving a shell. Alibaba's own DashScope route is $0.30/$1.50, so where you buy it changes the price nearly threefold.

#5
Claude Opus 4.8
Anthropic
Highest benchmark scoreBoth suites · 86.8
In: $5/1M
Out: $25/1M
Ctx: 1.0M

The ceiling, at 88.6 SWE-bench Verified and 84.6 Terminal-Bench 2.1 — the highest pair on this page — for $5/$25 per 1M. On a page weighted this hard toward price it still places, which is a real statement about the quality gap. Reserve it for the multi-file refactor that has already failed once on something cheaper, where an hour of your time costs more than the tokens. Opus 5 superseded it on July 24, 2026 at the identical $5/$25, so check which one your provider actually routes before assuming this row is what you are buying.

#6
Kimi K2.7 Code
Moonshot
Terminal-Bench only · 67.4
In: $0.95/1M
Out: $4/1M
Ctx: 262K

An open-weights 1T MoE (32B active) built specifically for coding, under a modified MIT licence, at $0.95/$4.00 with $0.19 cached input. It posts 67.4 on Terminal-Bench 2.1 — better than DeepSeek V4-Pro — and 44.7 on Terminal-Bench Hard, which makes it the strongest shell-loop model on this page under $1 input. The caveat is disclosure rather than capability: Moonshot's headline claims (+21.8% on its own Kimi Code Bench v2, roughly 30% fewer reasoning tokens than K2.6) are measured on a proprietary benchmark with no third-party replication. Note also that the older Kimi K2 Thinking models were discontinued on May 25, 2026; K2.7 Code and K2.6 are the live SKUs.

#7
GLM-4.7
Zhipu
Terminal-Bench only · 45.3Non-default tier
In: $0.6/1M
Out: $2.2/1M
Ctx: 200K

A workhorse mid-tier at $0.60/$2.20 with $0.11 cached input, sitting between the DeepSeek tiers and Kimi on price, and reachable from Z.ai direct, Bedrock, Vertex, Cerebras, Together and OpenRouter — useful in Kilo Code specifically because route choice is a config line, not a migration. It scores 45.3 on Terminal-Bench 2.1 in a reasoning configuration Artificial Analysis does not treat as its default tier, so read that number as an upper bound on what a plain API call buys. Its 200K context is exactly the floor of this page: enough for a working set, not enough to stop thinking about it. Zhipu also publishes a genuinely free GLM-4.7-Flash tier, which is a reasonable thing to point the read-only ask agent at.

#8
Gemini 3 Flash
Google
Longest contextNo published score
In: $0.5/1M
Out: $3/1M
Ctx: 1.0M

Listed on price, context and availability rather than on a score: Google publishes no SWE-bench Verified or Terminal-Bench 2.1 figure for it, and the Artificial Analysis entry that used to sit against this row was measured on a raised reasoning tier, so we removed it rather than pass it off as a default-tier result. What is verifiable is the shape of the deal — $0.50/$3.00 per 1M over a 1,048,576-token context, batch at half that, and context caching at $0.05/1M read. The trap for a Kilo Code user is that thinking is on by default at a high level and thinking tokens bill as output, so real sessions land well above the $3.00 sticker. Still preview-only as gemini-3-flash-preview.

How We Ranked These Models

Order is computed, not hand-picked. Each model is scored on the published pass-rate suites it has (SWE-bench Verified 0.55, Terminal-Bench 2.1 0.45), renormalised over the suites present, then charged 6 quality points for every doubling of its blended price (3:1 input-to-output tokens). Models with no published score are never given one: they are listed after every scored model, ordered by price, and labelled as unscored. Kilo Code runs the same kilo.json config in VS Code, JetBrains and a terminal CLI, so a model has to hold up in an editor loop and a shell loop rather than one or the other: SWE-bench Verified 0.55, Terminal-Bench 2.1 0.45. It is bring-your-own-key at zero markup and its users are the second-largest token consumers on OpenRouter, so price presses harder here than on any other page — every halving of the blended price is worth 6 points of pass rate. The 200K context floor keeps out models that cannot hold a working set without leaning on the middle-out transform. See the full ranking method.

Cost at Kilo Code volume
Kilo Code takes zero markup, so the provider's rate is your rate, and its users are the second-largest token consumers on OpenRouter. This page prices a halving of blended cost at six points of measured pass rate — the hardest price pressure of any ranking on the site.
Editor loop and shell loop together
One kilo.json config drives VS Code, JetBrains and the CLI, so a model has to apply patches that actually build and drive a terminal without losing the thread. SWE-bench Verified is weighted 0.55 and Terminal-Bench 2.1 0.45, rather than specialising in either.
Context the working set fits in
Models under 200K context are excluded from this page. Kilo Code's middle-out transform can compress a prompt that overflows, but a model that triggers it on ordinary tasks is quietly discarding the context the agent just gathered.
Cache economics, not just sticker price
An agent loop re-sends the same file tree every turn, so the cached-input rate often decides the bill rather than the headline input rate. Where a provider publishes one, it is named in the notes for that model.

Frequently Asked Questions

The same models, weighted for another tool

Each page states its own weighting, so the order changes with the workload.