Gemini 2.5 doesn't shut down on October 16. Vertex retires it on October 20, the Gemini API took its date down in August, and the cheapest way off 2.5 Flash-Lite still costs 2.86x.
If you read October 16 somewhere and blocked out next week for a migration, check which Google console you bill through first. One of them has a date. The other quietly took its date down in August.

Photo by Marcel Strauß on Unsplash
Which date applies to you
| You call 2.5 through | What Google says today |
|---|---|
| Vertex AI / Agent Platform | Retired October 20, 2026. Pro, Flash and Flash-Lite. |
| Gemini API, project that already used 2.5 | No shutdown date. "Not deprecated", served until further notice. |
| Gemini API, new project | Already blocked. 404, "no longer available to new users". |
Gemini API deprecations page (last updated October 9) and Google Cloud model versions page (updated October 10), both read October 11, 2026.
So Vertex customers have nine days. Everyone else on the Gemini API with an existing 2.5 workload can keep it running, and probably should while 3.8 Flash is on introductory pricing, because the cheapest like-for-like move off 2.5 Flash-Lite still nearly triples the bill.
Where October 16 came from
It was real. Archived copies of Google's Gemini API deprecations page from May through August 2 list October 16, 2026 against all three 2.5 models. The suggested replacement for 2.5 Flash drifted over those weeks, from 3 Flash Preview to 3.5 Flash to 3.6 Flash. Then the August 4 capture, with the page stamped "last updated 2026-08-03", shows "No shutdown date announced" in every row. No changelog entry mentions the change.
Articles written before August still carry the old date, and it was still showing up in search results when we looked. Kingy.ai withdrew it on August 23. Our own Gemini 2.5 Flash-Lite page dropped it the next day.
What replaced the date is an access rule. The deprecations page now says Google is limiting 2.5 "to users who have actively used them in the past" and points new projects at 3.5 Flash-Lite or 3.8 Flash. That has been live in practice since July. A developer on Google's forum reported 2.5 Pro returning "no longer available to new users" on July 28, and Google staff confirmed the policy two days later.
Vertex still has a date: October 20
Google Cloud's model versions page lists October 20, 2026 for gemini-2.5-pro, gemini-2.5-flash and gemini-2.5-flash-lite, and each model page carries a banner saying it is being retired that day. A footnote adds that the models "may remain accessible through the Gemini API" after the date. So the same weights, sold by the same company, retire on one platform and stay up on the other.
Someone asked about exactly this mismatch on the Gemini forum on August 11. Nobody from Google has answered.
Vertex names different successors from the Gemini API, too:
- 2.5 Pro: 3.8 Flash or 3.5 Flash
- 2.5 Flash: 3.8 Flash, 3.5 Flash-Lite or 3.1 Flash-Lite
- 2.5 Flash-Lite: 3.8 Flash, 3.1 Flash-Lite or Gemma 4
Neither platform names a Pro model as the successor to 2.5 Pro. 3.1 Pro is still a preview. And 3.5 Flash, one of the two picks for 2.5 Pro on Vertex, was deprecated on the Gemini API on October 8, with requests now routed to 3.6 Flash. Vertex retires 3.6 Flash on November 19. We wouldn't migrate to either.
The bill, before and after
A month of 100M input tokens and 10M output, the input-heavy shape most Flash and Flash-Lite workloads have. Standard paid tier, no caching, no batch. Old models in grey.
Our arithmetic from Google's Gemini API pricing page, read October 11, 2026. Rates are input / output per 1M tokens. 2.5 Pro and 3.1 Pro bill more above 200K-token prompts; this assumes shorter ones.
Read it by row. 2.5 Flash-Lite users have the worst of it: the cheapest Gemini replacement Vertex names, 3.1 Flash-Lite, is 2.86x the bill, and it has its own earliest shutdown date of May 7, 2027. (Vertex also lists Gemma 4, which we couldn't price.) The one the Gemini API tells new projects to use, 3.5 Flash-Lite, is nearly four times it. And 3.8 Flash turns a $14 month into $112.50 now and $225 from January.
2.5 Flash users get a quiet break. 3.5 Flash-Lite has the exact same $0.30 / $2.50 card, and 3.1 Flash-Lite is 27% cheaper. Whether a Lite model does your job as well is a question for your evals, but it's a cheaper question than jumping to 3.8 Flash, which roughly doubles the bill now and quadruples it in January.
Pro is the odd one out. 3.8 Flash halves its bill until December 31, then lands on $225 again, the same as 2.5 Pro on this mix. That match is a coincidence of the 10:1 ratio. Shift toward output and 3.8 Flash comes out cheaper; shift toward input and it costs more.
Thinking is on now, and you pay for it as output
Those numbers are a floor. Google's thinking docs list 2.5 Flash-Lite with thinking off by default. The newer models in that table all think by default: 3.5 Flash-Lite starts at minimal, 3.8 Flash at medium, and 3.1 Pro Preview at high. Thinking tokens bill at the output rate.
3.8 Flash and 3.1 Pro Preview can't switch it off at all. Their lowest setting is low. Google's latest-model guide warns that 3.8 Flash "can use more tokens on longer running and complex tasks" and suggests lowering the effort level or staying on 3.6 Flash, which Vertex is retiring in November.
There are code changes as well. The same guide's migration checklist says to move from thinking_budget to thinking_level (the old field still works for now), lists temperature, top_p and top_k as deprecated, and says candidate_count isn't supported. Gemini 3 Pro and 3 Flash don't do image segmentation, and Google's advice is to stay on 2.5 Flash with thinking off if you need it. On Vertex after October 20 that advice stops being possible to follow.
Search grounding loses most of its free allowance
This one is easy to miss because it isn't a token price. On the 2.5 models, Grounding with Google Search gives 1,500 free requests a day, then $35 per 1,000. On Gemini 3 and newer it's 5,000 free a month, shared across all those models, then $14 per 1,000.
Take an app doing 1,000 grounded requests a day. On 2.5 Flash every one of them fits under the daily allowance and grounding costs $0. On 3.8 Flash a 30-day month is 30,000 requests, 25,000 of them billable, so $350 a month. That's more than any token line in the chart above. And it's a floor: Gemini 3 bills each search query, and one request can fire several, where 2.5 billed per grounded prompt. By our count the lower rate only wins above roughly 2,400 grounded requests a day, and that's with one query each.
What we'd do this week
On Vertex, there's no avoiding it. Move 2.5 Flash traffic to 3.1 Flash-Lite or 3.5 Flash-Lite first and see if quality holds, since both are at or under today's price. Send only the work that fails there to 3.8 Flash, and set the thinking level explicitly so medium isn't picked for you. Don't plan on falling back to 2.5 through the Gemini API either. Unless that project has called 2.5 before, it will most likely count as a new user and get the 404.
On the Gemini API, with a project that already uses 2.5, we'd leave it alone for now. Nothing breaks on October 16. Spinning up a new project, though, gets you a 404, so keep the key you have.
Either way, put January 1 in the calendar. That's when 3.8 Flash's introductory price ends and every line of it doubles. You can run your own token counts through the calculator before then.
Google's pages, and the forum threads
- Gemini API: deprecations - no shutdown date for 2.5, limited-access note, new-project guidance
- Wayback Machine, August 2 capture and August 4 capture - the October 16 date, then its removal
- Google Cloud: model versions and lifecycle - October 20 retirement, Vertex replacements, 3.6 Flash retirement
- Gemini API: pricing - every token rate, 3.8 Flash introductory schedule, grounding
- Gemini API: thinking - default thinking levels per model
- Gemini API: changelog - October 8 deprecation of 3.5 Flash and 3.7 Flash
- Gemini API: latest model guide - 3.8 Flash token warning, migration checklist
- Gemini API: Gemini 3 developer guide - image segmentation
- Google AI forum: "no longer available to new users" and API vs Vertex date question