Skip to main content
TokenCost logoTokenCost
Model ReleaseSeptember 3, 2026ยท12 min read

Gemini 3.8 Flash arrives on an introductory price that is not counted from its introduction. It expires December 31 like the other two models on it, so the newest Flash gets 121 days of the discount and the one before it got 141.

Google shipped its third Flash model in six weeks on September 2, at $0.75 per million input tokens and $3.75 per million output. Those are the numbers Gemini 3.7 Flash and Gemini 3.6 Flash already carry, on every one of the thirteen priced rows across four service tiers. The interesting part is not that three generations cost the same. It is that the word Google attaches to the price is introductory, and the date it expires has nothing to do with when any of them were introduced.

Three lit vending machines at night against a dark wall, the third narrower than the other two

Photo by James Butterly on Unsplash

Six weeks, three models, one date

Gemini 3.8 Flash landed September 2 at $0.75 and $3.75 per million tokens, and there is nothing to compare it against on price because Gemini 3.7 Flash and Gemini 3.6 Flash print the same figures on all thirteen priced rows, across Standard, Batch, Flex and Priority. What is worth reading closely is the expiry, which belongs to the calendar rather than to any model. All three revert together on January 1, 2027, so the run each one gets depends on when it happened to launch: 141 days for 3.7 Flash, 121 for the newest. Every one of the thirteen rows doubles by exactly 2.0x, with no rounding residue, which makes the January arithmetic trivial and the escape routes nonexistent.

Three other things fall out of that. Gemini 3.6 Flash was not launched on this price at all; it shipped July 21 at $1.50 and $7.50 and spent 23 days there before Google swept it onto a discount built for a newer model. Gemini 3.5 Flash, the one that looks expensive, is simply the only current Flash at list: $132.00 against $60.00 on the job we price below today, and $132.00 against $120.00 once the reversion lands. And Google's own launch post recommends staying on 3.7 Flash for efficiency-first work, because 3.8 Flash might use more tokens, which at an identical rate card is the only way these two models can differ in cost at all.

Thirteen rows, three models, nothing to compare

RowThrough Dec 31From Jan 1
Standard input$0.75$1.50
Standard output, thinking included$3.75$7.50
Standard cached input$0.075$0.15
Standard cache storage, per hour$0.50$1.00
Batch input$0.375$0.75
Batch output$1.875$3.75
Batch cached input$0.0375$0.075
Flex input$0.375$0.75
Flex output$1.875$3.75
Flex cached input$0.0375$0.075
Priority input$1.35$2.70
Priority output$6.75$13.50
Priority cached input$0.135$0.27

Per million tokens except cache storage, which is per million tokens per hour. Read from Google's Gemini API pricing page on September 3, 2026, where the page footer says it was last updated the day before. This is one table for three models: Gemini 3.8 Flash, Gemini 3.7 Flash and Gemini 3.6 Flash all print these twenty-six figures, and we could not find a cell where they differ.

Google does not hide the second column in a footnote. Each price cell on the page carries two sentences, in the cell: $0.75 through December 31, 2026, then $1.50 starting January 1, 2027. That is better disclosure than most of the industry manages. It is also why the more interesting question is not what the number is but who the date belongs to.

An introductory price counted from nobody's introduction

Introductory pricing normally means a window attached to a product: ninety days from launch, six months from general availability, that shape. Google's Flash discount is attached to the calendar instead. December 31 is the same December 31 for a model that shipped in August and a model that shipped three weeks ago, which means the length of the introduction depends entirely on how late you were born.

ModelLaunchedOn $0.75 / $3.75 sinceDays of it
Gemini 3.6 FlashJuly 21, 2026August 13, 2026141
Gemini 3.7 FlashAugust 13, 2026August 13, 2026141
Gemini 3.8 FlashSeptember 2, 2026September 2, 2026121

Counting the launch day and December 31 at both ends. Gemini 3.8 Flash, the newest and the one Google is telling you to adopt, gets the shortest run at the discount of the three. Twenty days of release cadence cost it twenty days of cheap tokens.

Follow the cadence forward and the shape gets stranger. Google shipped these three 23 and 20 days apart, and describes it in its own words as the third Flash release in only six weeks. Another release on the same rhythm lands around September 23 and would inherit 100 days of discount. One in mid-October gets about 79. One in mid-December gets sixteen. A Gemini 3.9 Flash announced in the first week of January would launch at $1.50 and $7.50 and be described, accurately, as costing twice what its predecessor cost at launch.

We are not predicting those dates. The point is that under a calendar-anchored discount, a model's launch price is partly a function of when the release train happens to stop, and that is a weird thing for a rate card to encode.

3.6 Flash was moved onto this price, not born on it

The clearest evidence that the discount belongs to the date rather than to any model is Gemini 3.6 Flash, which spent its first twenty-three days outside it. It launched July 21 at $1.50 and $7.50 with no expiry note of any kind, and the Wayback Machine has Google's own pricing page on July 22 and again on August 12 showing exactly that. The next capture, August 13 at 18:11 UTC, shows $0.75 through December 31 and $1.50 starting January 1, plus the halved caching and storage lines.

August 13 is the day Gemini 3.7 Flash shipped. Google did not cut 3.6 Flash because 3.6 Flash needed a cut. It swept an existing model onto a promotion built for a newer one, on the newer one's launch day, and we found no changelog entry recording it. We wrote about the consequence at the time, which is that there is no longer a cost reason to prefer 3.7 Flash over 3.6 Flash. Three weeks later there is no cost reason to prefer either over 3.8 Flash, which is a strange place for a product line to be.

Worth knowing if you are checking our work against the press: VentureBeat's August 13 launch piece reported 3.6 Flash's price as $1.50 and $7.50 on the day it had already become $0.75 and $3.75. Google's archived page contradicts it. TrendingTopics got it right the following day, noting that the discount applied retroactively to a model that had launched at the higher prices in late July.

Every row doubles by exactly two, which is more restrictive than it sounds

Look down the two columns in the first table and divide. Input, 2.0. Output, 2.0. Cached input, 2.0. Cache storage, 2.0. Batch input, batch output, batch caching, flex, priority: 2.0 all the way down, thirteen for thirteen, with no rounding residue anywhere. We checked each pair rather than assuming, because a single row at 1.8 or 2.2 would change the advice in this section.

A uniform multiplier is convenient and merciless in the same breath. Convenient, because you do not need to model anything: whatever your January workload is, whatever the mix of fresh and cached input, however much thinking the model does, whichever tier you sit on, your Gemini Flash bill on January 1 is your December bill times two. There is no prompt shape that softens it.

Merciless for the same reason. Every lever you might pull was already available in December and already scales. Caching is 0.1x base input on both sides of the date, which is the multiplier Google holds on every Gemini Standard row we checked, 3.5 Flash and 3.5 Flash-Lite and 3.1 Pro included. It gets looser on the cheaper tiers, where rounding to the nearest cent pushes Flash-Lite's Batch and Flex cache reads to 0.133x, but on the Flash rows this post is about it holds. Batch and Flex are 0.5x on both sides. Priority is 1.8x on both sides. If you were not batching in December, batching in January takes you back to exactly your December standard-tier bill and no further: on the job we price below, that is $60.00 either way. The Batch API is worth precisely one January and nothing more.

That constant 0.1x cache multiplier is worth a note in its own right, because Anthropic broke its own version of it two days ago when Fable 5.1 shipped a cache read at 0.025x. Google has not. Its cache read is a tenth of input on every row we looked at, and it stays a tenth through the reversion.

Gemini 3.5 Flash is not the expensive one. It is the one at list.

Google still sells Gemini 3.5 Flash, now described on the models index as its legacy Flash model, providing baseline speed and foundational performance for routine, high-throughput workloads. Its status badge still says stable. It bills $1.50 and $9.00. Set next to 3.8 Flash it reads as an embarrassment: twice the input price and 2.4x the output price for a model four generations behind. That reading is an artifact of which model is on promotion.

We went back through Google's archived pricing pages from May 20 forward. Gemini 3.5 Flash has never carried a date, an asterisk, an introductory label or a reversion clause, on ai.google.dev or on Vertex. Its card in May is its card today. It is not the model whose price is unusual; it is the only current Flash whose price is a price.

Line3.5 Flash3.8 Flash nowGap3.8 Flash Jan 1Gap
Input$1.50$0.752.00x$1.501.00x
Output$9.00$3.752.40x$7.501.20x
Cached input$0.15$0.0752.00x$0.151.00x
Cache storage, per hour$1.00$0.502.00x$1.001.00x
Batch input$0.75$0.3752.00x$0.751.00x
Batch output$4.50$1.8752.40x$3.751.20x
Priority input$2.70$1.352.00x$2.701.00x
Priority output$16.20$6.752.40x$13.501.20x

On January 1 the input prices become the same number. Not close, the same: $1.50 against $1.50, on Standard, Batch, Flex and Priority alike. The whole surviving difference between a May model and a September model is $9.00 against $7.50 on output, which is 1.20x.

There is a small tell in the table that we think confirms the reading. Gemini 3.5 Flash's cache storage is $1.00 per million tokens per hour, and the newer models' is $0.50 rising to $1.00. The older model is already sitting on the post-reversion number. The newer ones are visiting.

One oddity while we were transcribing: on 3.5 Flash's Flex row, Google prints cached input as $0.08 where its own Batch row for the same model says $0.075. Both tiers are otherwise identical halves of Standard. We think that is a rounding slip in Google's table rather than a real four-cent-per-hundred-million difference, but it is the page's text, not ours.

One job, priced on both sides of New Year

A thousand agent turns, each carrying 40,000 tokens of context and producing 8,000 tokens out. Forty million input, eight million output. Small enough to be a week of one team's work, large enough that the second column matters.

Model and tierThrough Dec 31From Jan 1
3.8 / 3.7 / 3.6 Flash, Standard$60.00$120.00
3.8 Flash, Batch or Flex$30.00$60.00
3.8 Flash, Standard, 90% cache hits$35.70$71.40
3.8 Flash, Priority$108.00$216.00
3.5 Flash, Standard$132.00$132.00
3.5 Flash, Batch or Flex$66.00$66.00
3.5 Flash-Lite, Standard$32.00$32.00
3.1 Pro Preview, Standard, under 200k$176.00$176.00

Three things in that table are worth stopping on. The cached row doubles like everything else, $35.70 to $71.40, because a 90% cache hit rate does not change a uniform multiplier. Gemini 3.8 Flash on the Priority tier costs $108.00 today, which is 18.2% less than Gemini 3.5 Flash costs on Standard: the newest model at its most expensive setting undercuts the old one at its ordinary setting. On January 1 that flips to 1.64x and the arbitrage disappears.

And Gemini 3.5 Flash-Lite, which carries no expiry either, does not move at all: $32.00 in September and $32.00 in January. Today that is a little over half what 3.8 Flash costs. In January it is closer to a quarter, and the cheap model got there without changing a cent. The two models nobody is writing about are the two whose prices you can plan around.

Google Cloud describes the same discount as a different mechanism

On ai.google.dev the discount is a lower list price. You are charged $0.75 and that is what the invoice says. On Vertex, the footnote behind the asterisk says something else: promotional pricing provided through 50% credits back on net spend on select models within a given period. Same headline figures, different plumbing. One is a price, the other is a rebate against spend, and those do not always land in the same accounting month or against the same budget line.

Vertex also splits each model into two labelled rows, one through December 31 and one starting January 1, which is a clearer layout than two sentences inside a cell. And it applies the usual 10% surcharge for non-global regions, so a pinned European or Asian endpoint pays $0.825 and $4.125 now, and $1.65 and $8.25 later. That surcharge is the quiet one: our 40-million-token job costs $66.00 on a non-global Vertex endpoint today, which is more than the $60.00 global price and still only half the $120.00 everyone pays in January.

One more difference between the two pages. Vertex's banner lists four things on the discount: 3.8 Flash, 3.7 Flash, 3.6 Flash, and CodeMender using those models. The AI Studio pricing page does not mention CodeMender at all. If you are budgeting for it, the expiry applies there too, and only one of Google's two pricing pages tells you so.

The launch post contains a sentence telling you to stay on 3.7 Flash

Buried in the announcement, after the benchmark claims, Google writes that 3.8 Flash works harder and that at times the model might use more tokens, and suggests developers either use lower effort levels or continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads. We do not remember the last time a vendor shipped a model and recommended the previous one in the same paragraph.

It matters here more than it usually would, because the rate cards are identical. When a new model is cheaper per token you can absorb some token inflation and still come out ahead. At $0.75 and $3.75 on both models, every extra token 3.8 Flash emits is an unhedged increase. Here is what output inflation costs on the same 40-million-token job, holding input fixed.

Output tokensThrough Dec 31From Jan 1Bill change
8.0M, as measured on 3.7 Flash$60.00$120.00baseline
8.8M, 10% more$63.00$126.00+5.0%
9.6M, 20% more$66.00$132.00+10.0%
10.4M, 30% more$69.00$138.00+15.0%
12.0M, 50% more$75.00$150.00+25.0%

On this input-heavy shape, ten percent more output costs five percent more money, because output is exactly half the bill at this 5:1 token ratio. Google publishes no figure for how much more 3.8 Flash talks, so we are not going to invent one. What we can say is that the question is answerable on your own traffic in an afternoon, by replaying a sample through both models and comparing token counts rather than benchmarks, and that it is the only comparison between these two models where the answer is not zero.

The same arithmetic, run against 3.5 Flash, is the clearest statement of what January does. Today, Gemini 3.8 Flash could emit 27.2 million output tokens on this job, 3.4 times what we budgeted, and still bill no more than Gemini 3.5 Flash does at 8 million. After the reversion, its entire allowance is 9.6 million, or 1.2 times. The margin for a chattier model shrinks by two thirds on New Year's Day.

A Pareto frontier with one axis left blank

Wednesday's announcement had a second half: Gemini 3.8 Flash Cyber, a vulnerability-detection and patching model gated behind something Google calls the Fairwind Program, open to trusted government authorities, critical infrastructure operators and software maintainers. It carries real numbers: 47.2% pass@1 on CWE-Bench as run by Collinear, more than 70% success on an internal twenty-language vulnerability benchmark, and a Chrome Security claim of 2.6 times more correct patches than much larger commercial models.

The sentence we keep returning to is Google's own framing of the CWE-Bench result: 47.2% against a leading frontier model at 47.8%, yet offered at a significantly lower cost, which Google says puts it on the Pareto frontier. That is a concession on capability and a claim on price. The model appears on no pricing page. It has no entry on the models page, no model ID in the announcement, no context window, and no listing on OpenRouter. We searched Google's pricing page for the word cyber and got zero hits.

A Pareto frontier is a two-dimensional claim. Google has published one dimension and asserted the other. We are not saying the price is bad, because we cannot see it. We are saying that losing a benchmark and winning on cost is only checkable by someone who knows the cost, and right now that is a set of vetted defenders under an application process. It is also the second gated cyber model in two days to be written about without a rate. OpenAI said on September 1 that Astra meets the Critical cybersecurity threshold under its Preparedness Framework, the first model it has designated at that level, and that post carries no model ID, tier or token price either. Astra is not released yet, so its silence is easier to excuse. Gemini 3.8 Flash Cyber is being given to defenders now.

The feeds carrying this price do not carry the date

OpenRouter listed Gemini 3.8 Flash within hours of launch, on September 2 at 15:14 UTC, and gets everything right that it models: six endpoints across Google AI Studio and Vertex, $0.375 and $1.875 on flex, $0.75 and $3.75 on standard, $1.35 and $6.75 on priority, 1,048,576 context, 65,536 max completion, reasoning billed at the completion rate. Those match Google's own figures exactly.

What it does not carry is any field for December 31. There is nowhere in the endpoint record to say this price ends, so anything reading that feed sees $0.75 as a fact about the model rather than a fact about the next 120 days. This is a schema problem more than an OpenRouter problem, and it is going to hit every aggregator, cost dashboard and internal spend model that treats a rate as a scalar. We store the reversion as a dated schedule for exactly this reason, so our Gemini 3.8 Flash page shows both prices and the day the first becomes the second.

If your forecast for 2027 was built by pulling current prices from an API, the Gemini Flash lines in it are half of what they should be. So are a few others: we counted eight published AI API price expiries last week, and this one now covers three models plus CodeMender on a single date.

What we would do with the 120 days that are left

Move to 3.8 Flash if the benchmarks help you, because it costs the same as what you are running and Google has stopped pretending otherwise. Then measure your token counts on it against 3.7 Flash for a week, since that is the only axis on which these two models can differ in cost, and Google has told you which direction to expect.

Then put January 1 in the budget at exactly double. Not approximately double, and not softened by batching or caching if you are already doing either. If you are not batching, you have one 2x lever in reserve that gets you back to today's number and no further, and it is worth holding until you need it rather than spending it now.

The last thing is the one nobody is doing: look again at Gemini 3.5 Flash-Lite and at 3.5 Flash. Neither has an expiry. Both look unattractive today for a reason that stops applying in four months. If a workload sits somewhere near the boundary between Flash and Flash-Lite quality, the comparison you ran in September will give you a different answer in January, and the model will not have changed at all.

You can run any of these against your own traffic on our calculator or read the full cards side by side on the pricing page.

Where every number here came from

  • Google: Gemini API pricing - Read September 3, 2026; the page itself says last updated 2026-09-02. Source of all thirteen rows for 3.8, 3.7 and 3.6 Flash and both columns of each, the full 3.5 Flash and 3.5 Flash-Lite cards, the 3.1 Pro Preview 200k tier, the two-sentence expiry wording inside each price cell, the $0.08 Flex caching cell on 3.5 Flash, and the fact that the word cyber does not appear on the page
  • Google: Gemini 3.8 Flash and 3.8 Flash Cyber - September 2, 2026. The $0.75 and $3.75 figures, the footnote stating the introductory price expires December 31, 2026 and that $1.50 and $7.50 apply from January 1, 2027, the third Flash release in six weeks framing, HLE-Verified 54.9%, the CWE-Bench 47.2% against 47.8% and the Pareto frontier claim, the 70% twenty-language vulnerability figure, the Chrome Security 2.6x patches claim, the Fairwind Program and its eligibility wording, and the sentence recommending 3.7 Flash for efficiency-first workloads because the new model might use more tokens
  • Google: Gemini 3.8 Flash model page - The model code gemini-3.8-flash, 1,048,576 input tokens, 65,536 output tokens, stable rather than preview, thinking supported at low, medium and high with minimal returning an error, and Batch, Flex and Priority all supported. No knowledge cutoff is published, which is why this post does not state one
  • Google Cloud: Vertex AI generative AI pricing - The banner naming 3.8, 3.7, 3.6 Flash and CodeMender, the split rows through and after December 31, the 10% non-global region surcharge and every figure derived from it, and the asterisk footnote describing the discount as 50% credits back on net spend rather than as a reduced rate
  • Wayback Machine: Google pricing page, August 12, 2026 and the August 13 capture - The two snapshots that bracket the 3.6 Flash change. The earlier one shows $1.50 and $7.50 with no expiry note; the later one, 38 hours on, shows $0.75 through December 31. Captures from May 20 and July 22 supply the 3.5 Flash and 3.6 Flash launch cards
  • OpenRouter: Gemini 3.8 Flash endpoints - Pulled September 3, 2026. Six Google-operated endpoints, the creation timestamp of 2026-09-02 15:14:16 UTC, per-token rates matching all three tiers, 1,048,576 context and 65,536 max completion, and the absence of any field expressing the December 31 expiry
  • OpenAI: path to Astra, critical capabilities and frontier safeguards - September 1, 2026. The sentence designating Astra the first model to meet the Critical cybersecurity threshold under OpenAI's Preparedness Framework, cited here only for the fact that it carries no published price, model ID or tier. Astra was first named on August 7 in a separate post that did not designate it; the model is still pre-release
  • TokenCost: Gemini 3.7 Flash pricing - Our August 14 post on the previous instance of this, when the price gap between 3.6 and 3.7 Flash closed to zero on launch day
  • What we could not establish. Gemini 3.8 Flash's knowledge cutoff is on none of Google's pages, so we do not state one. Nothing about 3.8 Flash Cyber is priced anywhere we could find, so the comparison in that section is against a blank. Gemini 3.5 Flash's launch date is widely reported as May 19, 2026 at Google I/O but we could not find a Google document stating it, and the earliest archived pricing page carrying the model is dated the next day, May 20, which is consistent but not a statement; nothing in this post depends on the date, only on the card, which is verified unchanged from May 23. Google publishes no number for how many more tokens 3.8 Flash uses than 3.7 Flash, so the inflation table is a sensitivity range rather than a measurement. DeepSWE v1.1, Vals Finance Agent V2 and Harvey's Legal Agent Benchmark appear in the launch post with claims but no scores. And the release-cadence projection is arithmetic on two intervals, not a roadmap: Google has said nothing about a fourth Flash