Google cut Gemini 3.6 Flash in half on Thursday and launched Gemini 3.7 Flash at the same two numbers. The 50% is real, the saving is zero, and every line on both cards doubles on January 1.
Gemini 3.7 Flash went generally available on August 13 at $0.75 input and $3.75 output per million tokens, and the write-ups converged on one sentence: half the price of 3.6 Flash. Google's own phrasing is narrower and more careful. It says half the original 3.6 Flash cost, and the word carrying that sentence is original. Two Wayback snapshots of Google's pricing page, 38 hours apart, settle what happened. On August 12 Gemini 3.6 Flash was $1.50 and $7.50, flat, with no introductory wording and no end date on the row. On August 13 the same row reads $0.75 and $3.75 through December 31, 2026. Google discounted the outgoing model by exactly half on the same day it shipped the new one, and then described the new one as the discount. Move a workload from 3.6 Flash to 3.7 Flash this morning and your invoice changes by nothing at all.

Image source: Google
Gemini Flash has had three prices in five weeks
Read this as a calendar rather than a rate card, because the number is stable and the date is not. There is one price live today, it applies to both Flash models equally, and it has twenty weeks left to run.
Jul 21 to Aug 12, 2026
$1.50 / $7.50
3.6 Flash only, flat, no expiry stated
Aug 13 to Dec 31, 2026
$0.75 / $3.75
3.6 and 3.7 Flash, identical cards
From Jan 1, 2027
$1.50 / $7.50
Both models, exactly 2x every line
Google's own model card prints both models at the same price
The cleanest evidence is not a third-party tracker or an archived snapshot. It is page 5 of the Gemini 3.7 Flash model card, published the same day as the launch blog, which carries a comparison table with price rows in it. Those rows list Gemini 3.7 Flash at $0.75 and $3.75, and Gemini 3.6 Flash at $0.75 and $3.75. The footnote under them reads that for 3.6 and 3.7 Flash the introductory price expires on December 31, 2026, after which $1.50 and $7.50 apply. Google published a blog post describing 3.7 Flash as half the cost of 3.6 Flash and a model card showing them at parity, on the same morning.
| Line item | 3.7 Flash today | 3.6 Flash today | Both, from Jan 1 |
|---|---|---|---|
| Input / 1M | $0.75 | $0.75 | $1.50 |
| Output / 1M, incl. thinking | $3.75 | $3.75 | $7.50 |
| Cached input / 1M | $0.075 | $0.075 | $0.15 |
| Cache storage / 1M / hour | $0.50 | $0.50 | $1.00 |
| Batch input / 1M | $0.375 | $0.375 | $0.75 |
| Batch output / 1M | $1.875 | $1.875 | $3.75 |
| Vertex Priority in / out | $1.35 / $6.75 | $1.35 / $6.75 | $2.70 / $13.50 |
Vertex says the quiet part on the page itself. Its banner reads that Gemini 3.7 Flash and Gemini 3.6 Flash are offered with introductory pricing of $0.75 and $3.75 through December 31, 2026, naming both models in the same sentence. If you only ever read the Vertex pricing page you would never have formed the impression that 3.7 Flash undercuts its predecessor, because Google does not claim that there.
One caveat we cannot resolve. Every introductory row on Vertex carries an asterisk, and one nearby footnote defines promotional pricing as 50% credits back on net spend rather than a reduced rate. The same asterisk glyph is reused for an unrelated footnote about non-global endpoints, so we cannot tell which one governs. The rate tables state $0.75 and $3.75 directly, and that is what we have used. If you are on Vertex and the distinction between a discounted rate and a rebate matters to your accounting, ask your account team rather than trusting the table.
The archive shows 3.6 Flash never had an introductory price
This is the part that turns a marketing quibble into a fact. The Wayback Machine holds consecutive captures of Google's Gemini API pricing page on August 12 at 04:23 UTC and August 13 at 18:11 UTC, which pins the change to a window of about 38 hours.
| Capture | What the 3.6 Flash row said | Expiry language |
|---|---|---|
| Jul 22, 02:40 UTC | $1.50 in, $7.50 out, $0.15 cached | None |
| Aug 12, 04:23 UTC | $1.50 in, $7.50 out, $0.15 cached | None |
| Aug 13, 18:11 UTC | $0.75 in, $3.75 out, $0.075 cached | Through December 31, 2026 |
So for the whole of its first three weeks, Gemini 3.6 Flash carried an unqualified price with no end date. Two independent Google documents corroborate it. The July 21 changelog entry announcing 3.6 Flash never uses the word introductory, and the first appearance of that word anywhere in the changelog is the August 13 entry for 3.7 Flash. And the Gemini 3.6 Flash model card from July has its own price row reading $1.50 and $7.50, unqualified, in a table where Google explicitly annotated Claude Sonnet 5 as carrying a temporary discount. Google knew how to mark a promotional price in July. It did not mark its own.
The consequence is that January 1 means two different things depending on the model. For 3.6 Flash it is a reversion to the price it charged from launch until Thursday. For 3.7 Flash it is a doubling to a price it has never charged, on a card that was never anything but introductory.
One loose end in that July model card is worth closing, because it looks like an error and is not. Google lists GPT-5.6 Luna there at $1.00 and $6.00, which is five times what OpenAI charges today. Google was right when it wrote it: Luna was $1.00 and $6.00 until OpenAI cut it 80% on July 30, nine days after the card went up. A competitor price table has a shelf life of about a week in this market, which is worth remembering before quoting any vendor's comparison of anyone else.
One workload, two dates, and the model it stops undercutting
Take a coding agent that consumes 30M input and emits 6M output tokens in a month, a 5:1 ratio that sits between chat and heavy tool-calling. At Gemini Flash's 5x output multiplier the two legs weigh exactly the same, which makes the arithmetic unusually legible: $22.50 of input and $22.50 of output.
| Model | In / Out per 1M | 30M + 6M bill |
|---|---|---|
| DeepSeek V4 Flash, today | $0.14 / $0.28 | $5.88 |
| DeepSeek V4 Flash, off-peak from Aug 16 | $0.22 / $0.66 | $10.56 |
| GPT-5.6 Luna | $0.20 / $1.20 | $13.20 |
| DeepSeek V4 Flash, peak from Aug 16 | $0.44 / $1.32 | $21.12 |
| Gemini 3.7 Flash, batch, today | $0.375 / $1.875 | $22.50 |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | $24.00 |
| Gemini 3.7 Flash, today | $0.75 / $3.75 | $45.00 |
| Claude Haiku 4.5 | $1.00 / $5.00 | $60.00 |
| Vertex Priority, 3.7 Flash, today | $1.35 / $6.75 | $81.00 |
| Gemini 3.7 Flash, from Jan 1 | $1.50 / $7.50 | $90.00 |
| Grok 4.6, prompts under 200K | $2.00 / $6.00 | $96.00 |
| Gemini 3.5 Flash | $1.50 / $9.00 | $99.00 |
| Claude Sonnet 5 | $2.00 / $10.00 | $120.00 |
| GPT-5.6 Terra | $2.00 / $12.00 | $132.00 |
The row that matters is Claude Haiku 4.5. Today Gemini 3.7 Flash undercuts it by 25%, at $45.00 against $60.00. On January 1 the same job costs half as much again on Gemini as it does on Haiku, because Haiku 4.5 does not move and Gemini does. Anthropic reinforced the point days before this launch by making Claude Sonnet 5's $2 and $10 introductory rate permanent and cancelling its own scheduled September 1 increase, which is the precise opposite move to the one Google made on Thursday. We wrote that increase up in July on the assumption it would happen. It will not.
GPT-5.6 Luna is the awkward comparison. It bills $13.20 for the same work, which is 29 cents on the dollar against Gemini 3.7 Flash today and under 15 cents once January arrives, and it scores well enough that the gap is not explained away by capability alone. Gemini 3.7 Flash is not competing at the bottom of the market and the introductory price does not put it there.
Every line doubles, so nothing on your bill softens it
We checked each published rate against its January replacement expecting the discount tiers to move unevenly, because they usually do. They do not. Input, output, cached input, cache storage, both batch legs and both Vertex Priority legs are all exactly 2x. The practical reading of that is short: caching still takes 90% off, batching still takes 50% off, and Priority still costs 1.8x, on both sides of December 31. Whatever optimisation you have already done carries over unchanged and buys you none of the increase back.
Output is the line to watch, because Google bills thinking tokens there. The pricing row is labelled output price including thinking tokens, and the thinking documentation states that response pricing is the sum of output tokens and thinking tokens. Gemini 3.7 Flash takes a thinking_level of low, medium or high, defaults to medium, and rejects minimal outright, which is a setting 3.6 Flash still accepts and a second small thing the new model does worse. So the headline $0.75 input rate is the least relevant number on the card for agentic work, and it is the one every write-up led with. Run at high and your bill is governed by $3.75, going to $7.50.
There is a second-order version of the same trap. In stateless mode Google requires that thought blocks be resent unchanged on the following turn, and those arrive back as billed input. A long reasoning chain is therefore charged once at the output rate and again, as context, on every subsequent turn of the conversation. Our prompt caching breakdown has the method for working out whether the $0.075 cached rate recovers that, and at a 90% discount it usually does.
The trackers are still quoting a 2x gap that closed on Thursday
Artificial Analysis publishes a blended cost per million tokens on a 7:2:1 mix of cache hits, fresh input and output. Its page gives Gemini 3.7 Flash $0.58 and Gemini 3.6 Flash $1.16, which reads as a clean halving and is where a lot of the secondary coverage is getting its confidence. Recompute both from the published rates and the discrepancy is obvious.
| Blended 7:2:1 | Rates used | Per 1M |
|---|---|---|
| 3.7 Flash, as published | $0.075 / $0.75 / $3.75 | $0.5775 |
| 3.6 Flash, today's real rates | $0.075 / $0.75 / $3.75 | $0.5775 |
| 3.6 Flash, pre-Aug 13 rates | $0.15 / $1.50 / $7.50 | $1.1550 |
| 3.7 Flash, from Jan 1 | $0.15 / $1.50 / $7.50 | $1.1550 |
The $1.16 attributed to 3.6 Flash reproduces exactly from its pre-Thursday card, to the fraction of a cent, which identifies it as stale rather than wrong at the time. Both models blend to $0.5775 today. And the figure Artificial Analysis is currently showing as 3.6 Flash's cost is, to the cent, what 3.7 Flash will cost on January 1. If a tracker that exists to price these models has not caught a 50% change to a flagship card after a day, assume your own dashboards have not either.
Google published a table where its new model wins nine rows of twenty
The model card benchmarks 3.7 Flash against 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra and Muse Spark 1.2 across twenty rows. Google bolds the leader in each, and the tally is more honest than launch tables usually are: 3.7 Flash leads nine, GPT-5.6 Terra leads six and also tops the Intelligence Index at 57 against 56, Claude Sonnet 5 leads two, Muse Spark 1.2 leads one and ties the index, and Gemini 3.6 Flash leads one.
| Benchmark | 3.7 Flash | 3.6 Flash | Move |
|---|---|---|---|
| DeepSWE v1.1 | 65.3% | 48.6% | +16.7 |
| AutomationBench (private) | 30.4% | 17.0% | +13.4 |
| OSWorld-2.0 | 47.9% | 33.8% | +14.1 |
| GDP.pdf | 34.0% | 22.0% | +12.0 |
| Terminal-bench 3.0 | 14.9% | 5.4% | +9.5 |
| FrontierCode 1.1 Main | 43.6% | 34.4% | +9.2 |
| Terminal-bench 2.1 | 85.8% | 78.0% | +7.8 |
| GDM-MRCR v2, 8-needle | 97.0% | 91.8% | +5.2 |
| Code Arena (Elo) | 1588 | 1538 | +50 |
| Intelligence Index | 56 | 52 | +4 |
| CharXiv, no tools | 84.5% | 85.2% | -0.7 |
| CharXiv, with tools | 88.7% | 89.4% | -0.7 |
Twelve representative rows are above, with both losses included. Across the full table 3.7 Flash beats its predecessor on eighteen rows and loses two, and both losses are CharXiv, with and without tools, by seven tenths of a point each. That is a regression Google could have left out and did not, which is worth crediting.
Set the two halves of this launch side by side and the accounting comes out lopsided. The capability improvement is genuine and unevenly distributed, worth 16.7 points on agentic software engineering and four on the composite index. The price improvement, for anyone already on 3.6 Flash, is nothing, and it expires anyway.
Two cautions on numbers circulating for this model. Google published no GPQA Diamond, AIME, SWE-Bench Verified or MMMU figure for 3.7 Flash, and the scores appearing in search results under those names belong to Gemini 3 Flash, a different and older model. Separately, Google's blog gives 3.6 Flash's DeepSWE baseline as 49.0% where the model card says 48.6%. We have used the model card.
The flat card is worth more this week than the discount is
Here is the feature nobody put in a headline. Gemini 3.7 Flash has one input price and one output price at any prompt length, across a 1,048,576-token window. No threshold, no step, no surcharge. That is increasingly unusual, and the models it competes with have spent the last month going the other way.
| Model | Prices on one card | What triggers the change |
|---|---|---|
| Gemini 3.7 Flash | 1 | Nothing. A date, not a token count |
| Grok 4.6 | 2 | 200K prompt tokens, repricing the whole request |
| GPT-5.6 Terra | 2 | A short/long split OpenAI does not define on the page |
| DeepSeek V4 Flash, from Aug 16 | 2 | The clock. 7 peak hours a day at 2x off-peak |
| Qwen3.7 Flash | 3 | Steps at 32K and again at 256K of input |
A flat card is worth real money on workloads with unpredictable prompt sizes, which is most agent work. We have written up what the alternative costs on Qwen3.7 Flash's 32K step and on GPT-5.6's long-context tier, and the pattern is the same both times: the surcharge reprices the entire request rather than the excess, so a single token over the line can double a bill. Gemini 3.7 Flash has no line to cross. It has a calendar entry instead, which is at least a thing you can plan around.
The timing sharpens it. DeepSeek's peak and off-peak card takes effect at 16:00 UTC on August 16, two days from now, and it is an increase rather than a discount scheme: V4 Flash output goes from $0.28 to between $0.66 and $1.32. We measured the headroom for that rise last week before the numbers were published. So inside one week the cheap end of the market got a temporary halving with an expiry attached and the cheapest model on it got a permanent multiplier. Only one of those is reversible by waiting.
Switch for the benchmarks, budget for January
Our read. Move to 3.7 Flash, and do it for the benchmarks rather than the price, because the price is not a reason. It is materially better at agentic coding, it costs precisely what you are already paying for 3.6 Flash, and the only row where the old model wins is worth seven tenths of a point on a chart-reading benchmark. There is no argument for staying.
The thing to be careful about is the forecast rather than the migration. Twenty weeks of introductory pricing remain, counting today, and anything past that in a spreadsheet needs the doubled card in it. That is not a prediction on our part; it is printed on Google's pricing page, its model card and its Vertex banner in three consistent places. Whether Google holds the line is a different question, and Anthropic just demonstrated that these dates can be cancelled. We would plan for $1.50 and $7.50 and treat a reprieve as upside.
The broader habit worth taking from this week: a vendor comparing its new model against its own old one is quoting a price it controls, and it can change that price the same morning. The check takes a minute. Pull up the rate card for the model being compared against and read the row, rather than the sentence describing it. You can line all of these up on the pricing page or run your own token mix through the cost calculator. The 5:1 ratio used above flatters models with cheap output, and moving it reorders that table more than you would expect.
The receipts, including the archived ones
- Google: Gemini API pricing - Every rate in this post for both Flash models: $0.75 input, $3.75 output including thinking tokens, $0.075 cached input, $0.50 per 1M tokens per hour of cache storage, $0.375 and $1.875 batch, all through December 31, 2026, and the doubled figures from January 1. Also the Gemini 3.5 Flash card at $1.50 and $9.00 and Flash-Lite at $0.30 and $2.50, neither of which carries introductory language. Grounding is 5,000 free search requests a month shared across the 3.x family, then $14 per 1,000
- Google DeepMind: Gemini 3.7 Flash model card (PDF) - Page 5 carries the twenty-row benchmark table and the price rows listing both 3.7 Flash and 3.6 Flash at $0.75 and $3.75, with the footnote stating the introductory price expires December 31, 2026. Source for every benchmark figure above, including the two CharXiv rows where 3.6 Flash wins and the DeepSWE baseline of 48.6% that Google's blog gives as 49.0%
- Google: Introducing Gemini 3.7 Flash - The launch post, the $0.75 and $3.75 figures, the expiry sentence, and the exact phrase half the original 3.6 Flash cost. Note that the widely repeated framing of a 50% introductory price cut is press wording rather than Google's. Also the hero image on this post
- Wayback Machine: Gemini API pricing, August 12, 2026 - The capture showing Gemini 3.6 Flash at a flat $1.50, $7.50 and $0.15 with no introductory wording and no end date, 38 hours before the change. The July 22 capture is identical. This is what establishes that 3.6 Flash held unqualified pricing for its entire first three weeks
- Google: Gemini API changelog - The August 13 GA entry for 3.7 Flash, which is the first use of the word introductory anywhere in the changelog, and the July 21 entry for 3.6 Flash, which does not use it. There is no changelog entry recording the 3.6 Flash price cut at all. The July 21 entry also deprecates temperature, top_p and top_k on these models
- Google Cloud: Agent Platform pricing - The banner naming both 3.7 Flash and 3.6 Flash as carrying the same introductory rate, and the Priority tier at $1.35 and $6.75 rising to $2.70 and $13.50. This is also where the unresolved asterisk sits: a nearby footnote defines promotional pricing as 50% credits back on net spend, and the same glyph is reused for an unrelated endpoint footnote, so we cannot tell whether the Vertex introductory rate is a discounted rate or a rebate
- Artificial Analysis: Gemini 3.7 Flash - Intelligence Index 56 against 52 for 3.6 Flash, and the blended 7:2:1 costs of $0.58 and $1.16. The $1.16 reproduces exactly from 3.6 Flash's pre-August 13 rates and not from its current ones, which is what identifies it as stale
- Google: thinking - That response pricing is the sum of output and thinking tokens, the thinking_level values of low, medium and high with minimal rejected, and the requirement to resend thought blocks unchanged in stateless mode, which is what puts reasoning back on the input line on later turns
- Anthropic: pricing - Claude Haiku 4.5 at $1 and $5, Claude Sonnet 5 at $2 and $10, and the note stating that Sonnet 5's introductory rate is now standard and the September 1 increase to $3 and $15 will not occur. Anthropic's docs carry no date stamp for that reversal; secondary sources put it at August 10, which we have not been able to confirm on a primary source and have therefore not asserted
- OpenAI: API pricing - GPT-5.6 Luna at $0.20 and $1.20 and Terra at $2.00 and $12.00, both short-context. The page splits every model into short and long context columns without defining the threshold anywhere on it, which is why the comparison table above uses the short-context rate and says so
- DeepSeek: pricing - V4 Flash at $0.14 and $0.28 today, moving at 16:00 UTC on August 16 to $0.22 and $0.66 off-peak and $0.44 and $1.32 at peak, with peak defined as 01:00 to 04:00 and 06:00 to 10:00 UTC
- xAI: Grok 4.6 - $2 and $6 under 200K prompt tokens, $4 and $12 at or above it, and the rule that requests reaching 200K are billed at the higher rate for all tokens in the request rather than the excess
- Two gaps worth naming. Neither the pricing page nor the models page gives a knowledge cutoff; both model cards do, at March 2026, with the caveat that some domains stay limited to January 2025 in line with the wider Gemini 3 family. And there is no published RPM, TPM or RPD figure for either model, because Google's rate limits page now defers entirely to an account-gated dashboard. Batch enqueued tokens are the exception: Gemini 3.7 Flash appears there only in the Tier 1 table, at 3,000,000, with no Tier 2 or Tier 3 figure published. The 1,000,000,000 Tier 3 allowance being quoted around belongs to 3.6 Flash