Google's Lite tier keeps getting less lite. A year of releases took Flash-Lite output from 40 cents to $2.50.
Gemini 3.5 Flash-Lite shipped on July 21 at $0.30 per million input tokens and $2.50 per million output. The model it replaces charged $0.25 and $1.50. That is 20% more on input and 66.7% more on output, applied identically across all four service tiers, and it is the third consecutive Flash-Lite release to raise the price. The usual defense is that a smarter model finishes in fewer tokens. Here the efficiency gain is real and measurable, and it still does not get you back to even: Artificial Analysis ran its Intelligence Index on both models, used 21.8% fewer output tokens on the new one, and paid 63.3% more. One line on the rate card did move down, though, and if you are running audio it changes the answer completely.

Photo by Adhitya Sibikumar on Unsplash
Four Flash-Lites, four rate cards, one direction
Gemini 2.5 Flash-Lite was released on July 22, 2025. Gemini 3.5 Flash-Lite arrived on July 21, 2026, three hundred and sixty-four days later. Line up every Flash-Lite Google has shipped, with the shutdown dates from its deprecation page, and the trend is not subtle: three consecutive releases, three consecutive increases.
| Model | Released | Input / 1M | Output / 1M | Shutdown date |
|---|---|---|---|---|
| Gemini 2.0 Flash-Lite | Feb 25, 2025 | $0.075 | $0.30 | Retired Jun 1, 2026 |
| Gemini 2.5 Flash-Lite | Jul 22, 2025 | $0.10 | $0.40 | Oct 16, 2026 |
| Gemini 3.1 Flash-Lite | May 7, 2026 | $0.25 | $1.50 | May 7, 2027 (earliest) |
| Gemini 3.5 Flash-Lite (new) | Jul 21, 2026 | $0.30 | $2.50 | None announced |
Input tripled over the year. Output went from 40 cents to $2.50, which is to say the cheap tier now charges more for a million output tokens than it charged for six million a year ago. The cleanest way to see how far the tier has drifted is to notice what $0.30 and $2.50 already buys elsewhere on Google's own price list: those are the exact rates for Gemini 2.5 Flash, input and output both. The 2026 Lite model is priced to the cent like a full Flash model from the generation before it. The name did not move. The rung did.
The right-hand column is the part that makes this more than a grumble. Gemini 2.5 Flash-Lite, the last genuinely cheap option at $0.10 and $0.40, shuts down on October 16, 2026, which is 82 days from today. Its recommended replacement is 3.1 Flash-Lite, which carries its own earliest-possible shutdown of May 7, 2027. Google is careful to say those dates are the earliest a model might be retired rather than a commitment, so the runway may be longer. The direction is still one way.
Every way you can buy it moved by the same two multipliers
Google sells Gemini through four service tiers, and coverage of the launch quoted only the standard one. There is no tier where you escape the increase. Input is multiplied by 1.2 and output by 1.667 straight down the card, which tells you this was a deliberate repricing of the tier rather than a side effect of some new serving cost.
| Tier | 3.1 in / out | 3.5 in / out | Input change | Output change |
|---|---|---|---|---|
| Standard | $0.25 / $1.50 | $0.30 / $2.50 | +20% | +66.7% |
| Batch | $0.125 / $0.75 | $0.15 / $1.25 | +20% | +66.7% |
| Flex | $0.125 / $0.75 | $0.15 / $1.25 | +20% | +66.7% |
| Priority | $0.45 / $2.70 | $0.54 / $4.50 | +20% | +66.7% |
A separate change three weeks earlier makes the sticker price misleading for one group of users. On July 1, Google started charging a 10% premium on Vertex AI non-global endpoints for the generally available Gemini 3 and later families; before that date, regional endpoints were billed at global rates. If you are pinned to a specific region for data residency, 3.5 Flash-Lite costs you $0.33 and $2.75, not $0.30 and $2.50. To be precise about what this is and is not: the premium hits both models equally, so the migration itself is still the same +20% and +66.7%. What it does mean is that a team which was running 3.1 Flash-Lite on a regional endpoint in June, at global rates of $0.25 and $1.50, and moves to 3.5 Flash-Lite today, sees its input rise 32% and its output rise 83% from two unrelated changes landing three weeks apart. Neither number appears on the AI Studio page everyone quotes.
Fewer output tokens, bigger bill
Whenever a model gets more expensive per token, the reply is that it needs fewer of them. That defense is worth testing rather than accepting, and this is the rare case where somebody has run the controlled experiment. Artificial Analysis executes an identical evaluation suite against every model it tracks and publishes both the token count and the dollar cost, so the efficiency claim and the price increase land in the same number.
| Measure | 3.1 Flash-Lite | 3.5 Flash-Lite | Change |
|---|---|---|---|
| Output tokens used on the index | 55M | 43M | -21.8% |
| Cost to run the index | $93.72 | $153.08 | +63.3% |
| Blended price / 1M | $0.22 | $0.33 | +50% |
| Intelligence Index | 25 | 36 | +11 points |
| Output speed | 293.2 tok/s | 399.4 tok/s | +36% |
| Time to first token | 6.12s | 9.40s | +54% slower |
So the terseness is real. The new model genuinely says less, by about a fifth, and it is meaningfully smarter while doing it. It still cost 63.3% more to put through the same suite. On the output side alone the model would have needed to cut its token count by 40% to absorb the price rise, and it managed 21.8%.
Those two numbers do not multiply out to 63% on their own, and that is worth being straight about. A 21.8% token saving against a 66.7% output price rise gives roughly +30% if output were the only thing that changed. The measured bill came in at more than twice that, which means the extra is coming from the input side: back the list prices out of the totals and the 3.5 Flash-Lite run consumed several times the input tokens of the 3.1 run. That is what you would expect from a model doing more agentic turns per task. It is also the part nobody puts on a rate card, and it is why the output-price headline understates the increase rather than overstating it.
The last row deserves a note of its own. Flash-Lite is sold on latency, and time to first token got 54% worse. Thinking is on by default now, at the minimal level, and thinking tokens bill as output. Throughput went up and responsiveness went down, which is the wrong trade for anything sitting in front of a user.
We ran our own numbers on top of that, because an evaluation suite is not a production workload. Take $100 and spend it on a document-summarization call of 8,000 input tokens and 500 output tokens: 3.1 Flash-Lite gives you about 36,400 calls, 3.5 Flash-Lite gives you about 27,400, a quarter fewer. Now make it the agentic shape Google is actually marketing 3.5 Flash-Lite for, 2,000 tokens in and 4,000 out. The same $100 falls from roughly 15,400 calls to 9,400. The more you lean on the model the way Google wants you to, the harder the increase bites. Run your own shape through the cost calculator if your input-to-output ratio is nothing like either of those.
One line went the other way
Gemini 3.1 Flash-Lite charged a premium for audio input: $0.50 per million, double its text rate, with cached audio at $0.05. Gemini 3.5 Flash-Lite drops the separate audio line entirely. The pricing page now reads $0.30 for text, image, video and audio together, and OpenRouter's per-endpoint metadata corroborates it with an audio rate identical to the text rate across the standard, flex and priority tiers. Audio input got 40% cheaper in a release everyone filed as a price rise.
That is not a rounding detail for anyone doing transcription, call summarization or voice agents, because those workloads are enormously input-heavy. Feed in a million audio tokens and ask for a 20,000-token summary and you pay $0.35 on the new model against $0.53 on the old one, roughly 34% less. The crossover sits at the point where output tokens reach 20% of audio input tokens, and speech-to-summary work is nowhere near that line. If your pipeline is mostly listening, this release is a discount.
What the extra money actually buys
It would be easy to write this up as Google gouging the cheap seats, and that would be unfair. The model card puts the two price rows directly above the benchmark rows, in the same table, with 3.1 Flash-Lite in the adjacent column. Google did not bury the increase. It published it next to the reason for it, and the reason is substantial.
| Benchmark | 3.5 Flash-Lite | 3.1 Flash-Lite | GPT-5.4 mini | Haiku 4.5 |
|---|---|---|---|---|
| SWE-Bench Pro (Public) | 54.2% | 38.3% | 54.4% | 39.5% |
| Terminal-Bench 2.1 | 54.0% | 31.0% | 59.2% | 44.2% |
| GDPval-AA v2 (Elo) | 1140 | 642 | 1171 | 907 |
| OSWorld-Verified | 74.0% | 54.3% | 72.1% | 50.7% |
| MLE-Bench | 39.2% | 22.0% | n/a | n/a |
| CharXiv Reasoning (no tools) | 74.5% | 73.2% | 80.3% | 61.7% |
| GDM-MRCR v2, 128k avg | 72.2% | 60.1% | 42.7% | 35.3% |
A jump from 31% to 54% on Terminal-Bench 2.1, and from 642 to 1140 on GDPval-AA v2, is not a version bump. It lands within touching distance of GPT-5.4 mini on agentic coding while costing 60% less on input, and it beats both named rivals on long-context recall by a wide margin. If your workload was failing on 3.1 Flash-Lite, this is a different class of model and the higher price is straightforwardly worth it.
The CharXiv row is the exception and the one worth dwelling on: 74.5% against 73.2% is a 1.3-point gain, on chart reasoning, for 67% more per output token. Where the new model is not doing agentic work, the improvement thins out fast.
Two cautions on the table as a whole. The comparison is Google's own, against comparators Google chose. And the two model cards use almost disjoint benchmark sets, so there is no published head-to-head at all on the classic knowledge evals: no shared GPQA, MMMU, AIME or LiveCodeBench numbers exist between these two models. The improvement is documented for agentic, coding and long-context work specifically. Do not assume it transfers to classification or extraction, which is what most people actually run on a Lite model.
Google still points cost-sensitive users at the old model
The most telling thing about this launch is not on the pricing page. Five days after 3.5 Flash-Lite shipped, Google's own Gemini 3.5 migration guide still recommends the previous generation to anyone watching their bill. Under the heading for choosing a Flash model it names 3.1 Flash-Lite for low-cost, high-volume tasks and calls it a stable, long-term model optimized for efficiency. The migration checklist repeats the advice: if your use case is highly cost-sensitive, it says, consider migrating to 3.1 Flash-Lite instead. Gemini 3.5 Flash-Lite is not mentioned in either passage.
Developers noticed. The Hacker News thread on the launch drew 575 comments, and the pricing complaints are specific rather than reflexive. One commenter tracked the ladder back to 2.5 and called it boiling the frog, arriving independently at the same 6.25x figure. Another pointed out that the cheaper model you would fall back to now has a sunset date attached, so the price goes up and the alternative goes away together. That second point is the one that matters, because it is the difference between a price rise you can decline and one you cannot.
That reading is harsh but the rate card supports it. The only Gemini still priced where Flash-Lite was priced a year ago is 2.5 Flash-Lite, and it is 82 days from shutdown.
Four workloads, four different answers
This is not a post with a single verdict, because the rate card genuinely points different ways depending on what you send it. If you run audio, move to 3.5 Flash-Lite today; the 40% cut on audio input swamps everything else and you get the smarter model thrown in. If you run agentic or coding workloads that were struggling, the quality jump is large enough to justify the bill, and you should price it against GPT-5.4 mini rather than against the old Flash-Lite, because that is the tier it now competes in.
If you run high-volume text classification or extraction and 3.1 Flash-Lite was already clearing your quality bar, stay on it. It is stable, it is callable, and nothing in Google's documentation asks you to move. You have at minimum until May 2027, and Google's own guide endorses the choice. Just set a calendar reminder rather than trusting that the date holds.
And if you are on 2.5 Flash-Lite at $0.10 and $0.40, your October 16 deadline is the real story here, because there is no longer anything in the Gemini lineup at that price. Moving to the nearest replacement inside Google adds 150% to your input bill and 275% to your output bill on day one, before you have changed a line of code. That is the point at which it is worth looking outside Google entirely, since the budget tier elsewhere did not move the same way. Our pricing table has the current rates, and the comparison tool will put any two of them side by side.
The broader pattern is worth naming, because the industry story for three years has been that inference prices only fall. At the frontier that is still true. At the bottom of the range it has stopped being true, and Gemini Flash-Lite is the clearest example: same tier, same name, same job, 6.25 times the output price in a year. Cheap tokens are getting less cheap, and the model that replaces yours is the mechanism.
Sources
- Google: Gemini API pricing - Standard, batch, flex and priority rates for 3.5 and 3.1 Flash-Lite, plus the unified $0.30 audio input line
- Google DeepMind: Gemini 3.5 Flash-Lite model card - Benchmark table with price rows, 3.1 Flash-Lite comparison column, March 2026 knowledge cutoff
- Google: Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber - July 21, 2026 launch post and Google's own price-to-performance framing
- Artificial Analysis: Gemini 3.5 Flash-Lite - Intelligence Index 36, 43M output tokens, $153.08 to run the index, 9.40s time to first token
- Artificial Analysis: Gemini 3.1 Flash-Lite - Intelligence Index 25, 55M output tokens, $93.72 to run the index, 6.12s time to first token
- Google Cloud: Vertex AI generative AI pricing - Non-global endpoint rates and the July 1, 2026 effective date for the 10% regional premium
- Google: Gemini API deprecations - 2.5 Flash-Lite shutdown October 16, 2026; 3.1 Flash-Lite earliest shutdown May 7, 2027
- Google: What's new in Gemini 3.5 - Migration guidance still naming 3.1 Flash-Lite for cost-sensitive, high-volume work
- Google: Gemini thinking - Thinking on by default at minimal level for 3.5 Flash-Lite, billed as output tokens
- Ars Technica: Google reveals faster and cheaper Gemini 3.6 Flash - One of the few launch reports to note the Flash-Lite price is higher than 3.1
- Hacker News: Gemini 3.6 Flash and 3.5 Flash-Lite discussion - 575-comment launch thread and the developer reaction to the Flash-Lite ladder