Google cut Gemini Flash's output price and taught it to say less. Do both at once and the same answer costs about a third less.
Gemini 3.6 Flash shipped today, July 21, and its rate card is the rare kind that moves down. Input holds at $1.50 per million. Output drops from $9 to $7.50. On its own that is a 16.7% cut, but Google also says the model spends 17% fewer output tokens on the same task. Stack the two and the output side of an equivalent request runs roughly 31% cheaper than it did a day ago. The catch is the second model that launched alongside it: Gemini 3.5 Flash-Lite went the other direction and got more expensive.

Image source: Google
Two launches that point in opposite directions
Google shipped three models today. Gemini 3.6 Flash is the reasoning mid-tier, Gemini 3.5 Flash-Lite is the budget tier, and Gemini 3.5 Flash Cyber is a security model in limited pilot with no public price. The two you can buy tokens for tell opposite pricing stories, so put them side by side with the SKUs they succeed.
| Model | Input / 1M | Output / 1M | Cached in | Context |
|---|---|---|---|---|
| Gemini 3.6 Flash (new) | $1.50 | $7.50 | $0.15 | 1M, flat |
| Gemini 3.5 Flash (prev) | $1.50 | $9.00 | $0.15 | 1M, flat |
| Gemini 3.5 Flash-Lite (new) | $0.30 | $2.50 | $0.03 | 1M, flat |
| Gemini 3.1 Flash-Lite (prev) | $0.25 | $1.50 | $0.025 | 1M, flat |
| Gemini 3.6 Flash Batch | $0.75 | $3.75 | — | 1M, flat |
| Gemini 3.5 Flash-Lite Batch | $0.15 | $1.25 | — | 1M, flat |
The mid-tier moved down and the budget tier moved up. Gemini 3.6 Flash keeps 3.5 Flash's $1.50 input and trims output by a buck-fifty. Gemini 3.5 Flash-Lite jumped 20% on input and 67% on output over 3.1 Flash-Lite, from $0.25/$1.50 to $0.30/$2.50. If you were routing cheap traffic to Flash-Lite and assumed a Lite refresh would hold the line or drop, that assumption is now wrong, and your next invoice will show it.
One naming note, because Google's versioning is doing something odd here. The mid-tier is 3.6 while the budget tier is 3.5. They are not the same generation, and the Lite is a step behind the Flash on the version ladder even though they shipped the same morning. Read the number, not the launch date, when you reason about capability.
Why the output cut is bigger than it looks
A 16.7% output cut is a fine headline. It is not the whole number. Google's launch post says Gemini 3.6 Flash reduces output token usage by 17% versus 3.5 Flash on the same work, taking fewer reasoning steps and tool calls to land the same answer. Output pricing is a rate. Token count is a quantity. When both drop, they multiply.
Take a task that emitted 100,000 output tokens on 3.5 Flash. It cost 100K times $9 per million, so ninety cents. The same task on 3.6 Flash emits about 83,000 tokens at $7.50 per million, which is 62 cents. That is a 31% drop on the output line of one request, and none of it required touching your prompts. You just point at the new model ID.
The compounding only helps output-heavy work, since input held flat. A retrieval job that reads a 200K-token context and writes a 500-token answer barely moves. An agent that reasons through twelve tool calls and streams long completions is where the token cut earns its keep. If your traffic skews toward the second shape, the effective savings run closer to the 31% than the 16.7%.
Where 3.6 Flash lands against the cheap tier
Here is the part that complicates the good news. Even after the cut, Gemini 3.6 Flash is not the cheapest thing in its weight class on the rate card. Line up the mid and budget tiers for a 10M-input, 2M-output month and it sits mid-pack.
| Model | Input / 1M | Output / 1M | 10M / 2M month |
|---|---|---|---|
| DeepSeek V4 Flash | $0.14 | $0.28 | $1.96 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $8.00 |
| Claude Haiku 4.5 | $1.00 | $5.00 | $20.00 |
| GPT-5.6 Luna | $1.00 | $6.00 | $22.00 |
| Gemini 3.6 Flash | $1.50 | $7.50 | $30.00 |
| Grok 4.5 | $2.00 | $6.00 | $32.00 |
| Gemini 3.5 Flash (old) | $1.50 | $9.00 | $33.00 |
| GPT-5.6 Terra | $2.50 | $15.00 | $55.00 |
At $30 for the month, Gemini 3.6 Flash costs more per token than GPT-5.6 Luna ($22) and Claude Haiku 4.5 ($20), and it is roughly fifteen times DeepSeek V4 Flash's $1.96. The 17% token cut narrows that Luna gap in practice, since fewer emitted tokens partly offset the higher output rate, but it does not erase it. On a straight rate-card read, Google is not selling the cheapest mid-tier. It is selling the one that reasons and uses tools, and asking you to pay a small premium for the capability.
Gemini 3.5 Flash-Lite is the stronger cost story of the two. At $8 for the same month it undercuts every Western budget model on this list except DeepSeek, and the Lite now scores well enough on agentic work that it stops being a pure summarization toy. The catch is the one we already flagged: it costs more than the Lite it replaces, so the comparison that matters is not Flash-Lite versus Luna, it is Flash-Lite today versus Flash-Lite last week.
The benchmark sheet behind the prices
The prices only make sense next to the capability jumps Google published. Gemini 3.6 Flash posts a clear lift over 3.5 Flash on coding and computer-use benchmarks, which is the workload it is being priced for.
| Benchmark | 3.6 Flash | 3.5 Flash |
|---|---|---|
| DeepSWE (coding) | 49% | 37% |
| MLE-Bench | 63.9% | 49.7% |
| OSWorld-Verified (computer use) | 83.0% | 78.4% |
| GDPval-AA v2 (ELO) | 1421 | 1349 |
The MLE-Bench line is the one to watch. A 14-point gain on machine-learning engineering tasks between point releases is large, and it lands in the same window as the token-efficiency claim. The model is not just cheaper per token, it is getting more done per token on the exact agentic work Google keeps pointing at. On Artificial Analysis it also clocks the fastest output speed of any model they track, north of 300 tokens per second, which is the latency half of the pitch.
Flash-Lite's numbers explain its price increase. It goes from 31% to 54% on Terminal-Bench 2.1 over 3.1 Flash-Lite, and its GDPval-AA v2 score nearly doubles from 642 to 1140. Google did not raise the Lite price to pad margin. It raised it because the model crossed from cheap-and-limited into cheap-and-useful, and repositioned the rate card to match. Whether that trade is worth 20 to 67% more per token is a question only your eval can answer.
Cyber, the Pro delay, and the Gemini 4 tell
Two things in the launch matter for anyone planning a quarter of spend. First, Gemini 3.5 Pro is not here. Google confirmed it is still testing with partners and will ship broadly when ready, which is the second time the Pro tier has slipped. If your roadmap assumed a 3.5 Pro rate card this month, keep budgeting against 3.1 Pro or a competitor flagship for now.
Second, DeepMind said in the same post that it has started its most ambitious pre-training run yet, for Gemini 4. That is a signal, not a date, but it frames today's launch: Google is trimming and tuning the Flash line while the next generation cooks. The third model, Gemini 3.5 Flash Cyber, is a vulnerability-finding security model behind a limited pilot with no token price, so it does not enter the cost math yet. Worth watching if you run security tooling, irrelevant to your bill today.
So which Flash should you run
If you run agent or coding traffic on Gemini 3.5 Flash today, switching to 3.6 Flash is the easy call. Same input price, lower output price, fewer output tokens, and a better benchmark sheet. The only reason to hesitate is a workload that somehow got tuned around 3.5 Flash's exact output length, and that is rare. Flip the model ID, watch one billing cycle, keep the savings.
If you run cheap bulk traffic on Gemini 3.1 Flash-Lite, do not auto-upgrade to 3.5 Flash-Lite. The new Lite is better and more expensive, and for summarization, classification, or single-shot extraction you may be paying 20 to 67% more per token for capability those tasks never use. Run the two side by side on your real prompts before you migrate, and if the quality is a wash, 3.1 Flash-Lite is still on the menu at the old price.
And if you are shopping the cheap tier cold, Gemini 3.6 Flash is not the automatic pick. GPT-5.6 Luna and Claude Haiku 4.5 are cheaper on the rate card, and DeepSeek V4 Flash is in a different universe on price. Gemini earns the premium only where the reasoning, the computer-use scores, and the token efficiency show up in your own numbers. The pricing page has every current card side by side, and the calculator runs your exact token mix against all of them. For the model this one replaces, see the Gemini 3.5 Flash write-up, and for the budget tier's history see the 3.1 Flash-Lite piece.
Sources
- Google: Gemini 3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber - July 21, 2026 launch post, benchmark sheet, and 17% token-efficiency claim
- Google: Gemini API pricing - $1.50/$7.50 for 3.6 Flash, $0.30/$2.50 for 3.5 Flash-Lite, cached and batch tiers
- Google DeepMind: Gemini Flash model page - 1M context, 64K max output, reasoning model spec
- Artificial Analysis: Gemini 3.6 Flash - Intelligence Index, 300+ tok/s output speed, blended cost
- 9to5Google: Gemini 3.6 Flash launch - Launch coverage, knowledge cutoff, Pro delay, Gemini 4 pre-training note
- AI Pricing Guru: OpenAI GPT-5.6 pricing - GPT-5.6 Luna $1/$6, Terra $2.50/$15 for comparison table
- IntuitionLabs: AI API pricing comparison - Claude Haiku 4.5 $1/$5, Grok 4.5 $2/$6 cross-check