Skip to main content
TokenCost logoTokenCost
GuideAugust 10, 2026·13 min read

Imagen 4 stops answering on August 17. Google's pages send you to two different replacements and one of them dies six weeks later. The model that costs less than Imagen 4 and has no end date of its own is named in neither.

Next Monday three endpoints go dark: imagen-4.0-fast-generate-001, imagen-4.0-generate-001 and imagen-4.0-ultra-generate-001. The interesting part is not that a model is being retired. It is that the retirement changes the billing unit. Imagen 4 sold you an image for a flat price that did not care about resolution, and the replacement sells you tokens, so the same picture now costs more if it is bigger. For anyone on the Fast tier generating at 1K that is $0.02 becoming $0.067, a 3.36x increase. If you were taking 2K output, which Imagen 4 charged nothing extra for, it is 2.52x on the Standard tier and 5.04x on Fast, with a documentation caveat about Fast that we will get to. Ultra customers pay 12% more and will barely notice. Underneath that spread sit two findings we did not expect. Five Google pages disagree about where to migrate, splitting two ways over a 72% price difference, and the freshest of them points at the cheaper model, which itself dies on October 2. Meanwhile Gemini 3.1 Flash Lite Image, batched at $0.0168, is the only Google image model that both undercuts Imagen 4 Fast and carries no end date of its own, and it is named in none of the notices. Below: the full token grid, what a month of a thousand images a day does before and after, the two cost lines that did not exist on an Imagen bill, and the response shape that returns success with no picture in it.

Macro of a film negative with a hard edge where bright orange light leak falls off into black

Photo by Chase Yi on Unsplash

Seven days, and a second date that already went by

The sentence on Google's Imagen 4 model page is short and does not hedge:

"The Imagen 4 standard, ultra, and fast endpoints are deprecated and will be shut down on August 17, 2026; migrate to Gemini 3.1 Flash Image to avoid service interruptions."

That went up on June 15, 2026, which is 63 days of notice, and the changelog entry for the same day uses identical wording. Google draws a distinction between the two verbs it is using. A deprecation is the announcement that support is ending. A shutdown is when the model is, in its words, completely turned off and the endpoint is no longer available. Monday is the second thing.

One qualifier belongs here, because it cuts against reading the date too literally. The deprecations page describes its published dates as the earliest possible dates on which a model might be retired, and Firebase's migration guide phrases the same event as shut down as early as August 17, 2026. So August 17 is a floor. Plan against it anyway, but do not be surprised if the endpoint answers on the 18th, and do not treat that as permission.

The part that surprised us is that this is the second shutdown of the same three model IDs. Vertex AI announced its own deprecation on March 24, 2026 with a discontinuation date of June 30, 2026, covering the Imagen 4 IDs plus the entire Imagen 3 line and the older imagegeneration@002 through @006 family, along with imagetext@001. That date passed six weeks ago, and Vertex's generative AI pricing page has since dropped its Imagen rows entirely; it now presents as Agent Platform pricing and lists Gemini rates only. The Imagen 4 model card is still standing, though, still updated as recently as August 4, and still recommending gemini-2.5-flash-image as the replacement. So on Vertex the prices are gone while the migration advice remains, and the advice points at a model with 53 days left. We did not test whether the endpoints still answer, because that means a live billed call.

The old price did not know how big the image was

This is the fulcrum of the whole migration, and it is easy to miss because it is a property of the old rate card rather than the new one. Imagen 4 charged per image, full stop:

Imagen 4 tierPer imageNative resolutionsBatch discount
Fast$0.021K, plus 2K on one page of twoNone
Standard$0.041K and 2K, both pages agreeNone
Ultra$0.061K, plus 2K on one page of twoNone

Read the third column. The flat fee did not scale with pixels, so on Standard a 2048 by 2048 image cost the same four cents as a 1024 by 1024 one. That was free quality, and it is about to be repriced harder than anything else in this migration.

We have to be careful about which tiers got 2K, because Google's two relevant pages contradict each other and we are not going to pick the one that makes our headline bigger. The Vertex output-resolution page, updated August 4, 2026, lists 2K support for imagen-4.0-generate-001 and imagen-4.0-fast-generate-001, so Standard and Fast but not Ultra. The Gemini API Imagen guide, updated July 16, 2026, says the image-size parameter is supported only for the Standard and Ultra models, so Standard and Ultra but not Fast. The two pages agree on exactly one tier. Every 2K figure in this post is arithmetically correct given the token counts, but only the Standard row rests on undisputed documentation. If you are on Fast and requesting 2K today, your own logs are better evidence than either page, and they are worth ten minutes before Monday.

The fourth column matters too, in the other direction. Imagen 4 never had a batch tier. Google's Batch API documentation says the feature is only available with the generateContent API, and Imagen used :predict. The word imagen appears nowhere in that document. So the honest frame for this comparison is asymmetric on purpose: the old model had no discount available to it and the new one has a 50% one. We will use both columns throughout rather than picking whichever makes the headline bigger.

Now the price knows, and it charges by the pixel

Gemini 3.1 Flash Image bills image output at $60.00 per million tokens, halved to $30.00 through the Batch API. An image is a fixed token count that scales with resolution, so the flat fee becomes a ladder. Every figure in the last two columns is ours, computed from tokens times rate, and every one reproduces the per-image number Google publishes in its own footnote to the cent.

Output sizeTokensStandardBatch
0.5K (512 px)747$0.045$0.022
1K (1024 px)1,120$0.067$0.034
2K (2048 px)1,680$0.101$0.050
4K (4096 px)2,520$0.151$0.076

Put the two cards against each other and the migration is not one number, it is a twelve cell grid. Multipliers at standard pricing, old tier down the side and new resolution across the top:

Migrating fromTo 0.5KTo 1KTo 2KTo 4K
Fast, $0.022.24x3.36x5.04x7.56x
Standard, $0.041.12x1.68x2.52x3.78x
Ultra, $0.060.75x1.12x1.68x2.52x

Two readings fall straight out. Ultra customers have almost nothing to complain about, and Ultra users who drop to 0.5K actually save 25%. Fast customers absorb the entire cost of Google changing its mind about billing units, at 3.36x for the 1K output nobody disputes they had.

One trap in the source data worth naming, because it is easy to walk into. Gemini 3 Pro Image, the tier above the model in this table, consumes 1,120 tokens for a 2K image and 2,000 for a 4K one. Those figures are not transferable: on 3.1 Flash Image a 2K image is 1,680 tokens and a 4K image is 2,520, which are the counts published in Google's pricing footnote and the Vertex model documentation, and the only counts that reproduce the per-image prices when multiplied out. Cross-apply the Pro numbers to Flash and you will budget $0.067 for a 2K image and be invoiced $0.101, a 50% miss on the exact line item you were trying to control.

Somebody has already measured this from the outside, which is worth more than our arithmetic. A developer posted a side-by-side on Google's own developer forum on July 10 showing what the API actually reported per request. Their generateContent counts came back 747, 1,120, 1,680 and 2,520 across the four resolutions, matching the documented table exactly. The Interactions API, on the same prompts, reported 1,120 at every resolution, which if it had driven billing would have undercharged a 4K image by more than half. A Google representative confirmed on August 3 that the broken token counts had been patched, saying the API should now report the correct number for various image resolutions.

We flag that for two reasons. It is independent confirmation that the numbers in our table are the ones the model actually emits, which is a better foundation than a footnote. And it is a reminder that a usage field reporting the wrong count for three weeks is the kind of thing you only catch by reconciling reported tokens against the invoice, which almost nobody does on an image pipeline. The same thread notes that multi-turn editing at 2K and 4K was returning 404 on both APIs, still under investigation at the time of the reply, so if your migration plan involves iterative edits at high resolution that is worth testing before Monday rather than after.

Five Google pages, two destinations, a 72% spread

We went looking for a single authoritative migration target and found the question genuinely open. Five Google-owned pages split two ways, and the split does not follow how recently they were touched. Read on August 10, 2026, with the last-updated dates Google itself prints on them:

Google pageUpdatedSays migrate to1K image
DeprecationsAug 3, 2026Gemini 3.1 Flash Image$0.067
Imagen model pageJun 15, 2026Gemini 3.1 Flash Image$0.067
PricingAug 5, 2026Gemini 2.5 Flash Image$0.039
Imagen guideJul 16, 2026gemini-2.5-flash-image$0.039
Vertex Imagen 4 cardAug 4, 2026gemini-2.5-flash-image$0.039

The awkward part is that recency does not settle it. The two freshest pages of the five, the pricing page at August 5 and the Vertex Imagen 4 card at August 4, both point at Gemini 2.5 Flash Image. The deprecations table naming 3.1 Flash Image was touched on August 3, a day earlier, and the Imagen model page carrying the headline shutdown sentence has not been updated since June 15. So the newest advice on the newest pages is for the older model, and it is a genuinely tempting instruction: at $0.039 per image the 2.5 model is 42% cheaper than the one the deprecations page names, which makes the migration nearly free.

Then check the deprecations table for the model you just moved to. Gemini 2.5 Flash Image shuts down on October 2, 2026, 46 days after Imagen 4. Take the advice on the freshest pricing page and you migrate twice in seven weeks, and the second migration is the one you were being warned about in the first place. There is a further wrinkle that reads like a bug: the replacement the deprecations table lists for 2.5 Flash Image is gemini-3.1-flash-image-preview, and that model already shut down on June 25, 2026. The recommended destination is itself dead.

The cheapest way out has no end date and no mention

At standard latency, every migration pointer in that table sends Imagen 4 users to a model that costs more per image than what they were using. Batch changes that for one of them: Gemini 2.5 Flash Image batched is $0.0195, which does undercut Imagen 4 Fast's $0.020. But it dies on October 2, so that is a seven week reprieve rather than a migration, and it is worth naming precisely because it is the trap the freshest pages walk you into.

Apply both filters at once, cheaper than what you were paying and no announced end of life, and exactly one Google model qualifies. It was released six weeks ago and it is absent from all five pages above.

Gemini 3.1 Flash Lite Image went GA on June 30, 2026. It bills image output at $30.00 per million tokens, exactly half the Flash Image rate, and it produces 1K output only. At 1,120 tokens that is $0.0336 standard and $0.0168 batched. Line those up against what the Fast tier used to cost:

Option for a 1K imagePer imagevs Imagen 4 FastStatus
Flash Lite Image, batch$0.0170.84xGA, no shutdown date
Gemini 2.5 Flash Image, batch$0.01950.97xDead October 2
Imagen 4 Fast$0.0201.00xDead August 17
Flash Lite Image, standard$0.0341.68xGA, no shutdown date
Flash Image, batch$0.0341.68xGA, no shutdown date
Gemini 2.5 Flash Image$0.0391.94xDead October 2
Flash Image, standard$0.0673.36xGA, no shutdown date

The top row is the finding. Batched Flash Lite Image at $0.0168 per image is cheaper than Imagen 4 Fast has ever been, on a model with no announced end of life, and no Google page we could find suggests it to the people being evicted. There is a real tradeoff attached, which is the 24 hour turnaround window that batch imposes and the 1K resolution ceiling. Neither disqualifies the large asynchronous jobs that make up most bulk image generation.

Notice also that two different rows land on $0.034. Flash Lite at standard latency and Flash Image through the batch queue cost the same per picture, which turns the choice into a straight question about what you actually want: a better model on a 24 hour delay, or a lighter model right now, for identical money. That is a more useful way to spend the decision than arguing about which one Google meant.

A thousand images a day, before and after

Per-image prices in thousandths of a dollar are hard to feel. Here is the same set of choices at 1,000 images a day over a 30 day month, which is a modest product catalogue or a mid-sized avatar feature.

SetupMonthlyChange
Was: Imagen 4 Fast at 1K or 2K$600baseline
Flash Lite Image, batch$504-$96
Flash Lite Image, standard$1,008+$408
Flash Image at 1K, batch$1,008+$408
Flash Image at 1K, standard$2,016+$1,416
Flash Image at 2K, standard$3,024+$2,424

The spread between the best and worst honest reading of the same instruction is $504 a month against $3,024 a month, six times over, and both ends are Google models generating a picture from a prompt. Worth being precise about what the cheap row costs you in kind rather than in dollars: Flash Lite Image is 1K only and batch means a 24 hour window, so $504 buys smaller pictures on a delay. The bottom row buys 2K on demand. Everything between them is set by two config values nobody reviews, which Flash variant you name and what resolution you ask for, and neither is the model-selection decision people think they are making.

Line items an Imagen invoice never had

Switching to a general-purpose multimodal model means inheriting its billing surface. Three things appear that did not exist before, and one of them we cannot put a number on.

Input tokens. Text prompts now bill at $0.50 per million. Imagen 4 capped prompts at 480 tokens, so even a maximal one costs $0.00024, which is 0.36% of a 1K image. Ignore it. Reference images for editing bill at 1,120 tokens each, so a 14 image edit request adds about $0.008. Also small, but no longer zero.

Thinking tokens, and this is the honest gap in this post. Gemini 3 image models are thinking models. Google states that thinking is enabled by default and cannot be disabled through the API, and separately that thinking tokens are billed. They bill at the text rate of $3.00 per million rather than the $60.00 image rate, and the interim thought images are explicitly not charged. What Google does not publish anywhere we could find is how many thinking tokens a typical image request consumes. So we will not invent one. For scale: 500 thinking tokens adds $0.0015, about 2.2% on top of a 1K image, and 2,000 adds $0.006, about 8.9%. Treat every per-image figure in this post as a floor, and treat thinking_level: high as an unmeasured multiplier rather than a free quality dial.

Grounding. Search grounding is available on this model and priced at $14 per 1,000 requests once you pass 5,000 free requests a month, shared across all Gemini 3.x models. Imagen 4 had no grounding at all, so this is a cost line that cannot have regressed, only appeared. Worth knowing before someone enables it to improve prompt fidelity.

And one thing that is absent on both sides: there is no free tier for any of these models. Every Gemini 3.1 Flash Image row on the pricing page reads not available under Free Tier, and the same is true of Flash Lite Image, Pro Image, 2.5 Flash Image and Imagen 4 itself. If you were hoping to prototype the migration for nothing, you cannot, and you could not before either.

The response that succeeds without returning a picture

Most coverage of this migration describes it as a one-line change from generate_images() to generate_content(). Three things move, not one. The method changes. The REST endpoint changes from :predict to :generateContent, which is a different request body rather than a different path. And the response goes from a typed image object with a generated_images list to content parts you have to walk.

That third change carries the failure mode we would guard first. Gemini can return FinishReason.NO_IMAGE: a successful response, no exception raised, no picture inside. Code written against Imagen's typed response will read that as an empty list and carry on. Check the finish reason explicitly before you assume you have an image, and decide up front whether a NO_IMAGE response should be retried, because a retry is another billed request.

Several Imagen parameters have no Gemini equivalent at all, per Google's own migration guide: numberOfImages, which returned one to four images from a single call, negativePrompt, imageFormat (output is PNG only now), personGeneration and addWatermark, since SynthID is now always applied. The multi-image case is not strictly gone, just unparameterised: the model page publishes a 32,768 output token cap, which by our arithmetic is about 29 images at 1K. You get there by asking in the prompt instead of setting a field, which is harder to validate and harder to cost.

Google vacated the price bracket it used to own

At $0.04, Imagen 4 Standard undercut every rival flagship tier while sitting above the cheap tiers that exist to be cheap. Flash Image at 1K does not sit there any more. Prices below are per image at roughly 1024 square, read off each vendor's own docs on August 10, 2026, and they mix quality tiers on purpose so you can see where the brackets fall.

ModelPer image
OpenAI gpt-image-2, low$0.006
OpenAI gpt-image-1-mini, medium$0.011
Gemini 3.1 Flash Lite Image, batch$0.017
Imagen 4 Fast (dead Monday)$0.020
xAI grok-imagine-image$0.020
Imagen 4 Standard (dead Monday)$0.040
OpenAI gpt-image-1, medium$0.042
Ideogram 4.0 Default$0.060
Gemini 3.1 Flash Image at 1K$0.067

Post-migration, the model Google tells you to use is the most expensive row on that list. It sits above Ideogram 4.0 Default and above gpt-image-1 at medium quality, in a bracket where Imagen 4 Standard used to undercut both. We are not claiming those models are equivalent in output quality, and Flash Image is a newer and more capable system than most of what it now costs more than. Nor is the list exhaustive; we cut the FLUX rows because Black Forest Labs no longer publishes a per-image table we could read directly, and we are not going to price a competitor off somebody else's blog. The point is narrower: if price was your reason for being on Imagen 4, Google is no longer the cheap answer, and the deadline is a reasonable moment to check whether it still needs to be Google at all.

Before anyone frames this as a Google-specific mess: OpenAI removed DALL·E from its API on May 12, 2026, and has shutdown dates on the board for gpt-image-1 (October 23, 2026) plus gpt-image-1-mini, gpt-image-1.5 and chatgpt-image-latest on December 1, 2026. Four of its five image models are dated. xAI retired grok-imagine-image-pro in May. Image endpoints are the most churn-prone surface in this market, and pinning an exact model ID with a calendar reminder is worth more here than anywhere else in your stack.

It comes down to who is waiting

The migration decision is not really about which model. It is about two questions you can answer from your own logs in ten minutes: what resolution am I actually asking for, and does anyone wait for the result.

If nobody waits, meaning the images land in a catalogue, a digest, a nightly render, then batch is the answer and the model choice is almost secondary. Flash Lite Image batched at $0.0168 is cheaper than what you are paying now. Flash Image batched at $0.0336 adds about two thirds to the bill, which on a thousand images a day is $408 a month for a considerably better model. Both are survivable, and neither is what you land on by following the migration notice, which names a model and says nothing about the tier that halves it.

If somebody does wait, the resolution question decides it. Go and check what you are actually requesting, because Imagen 4 charged nothing extra for 2K and there is a decent chance a default was set once and never revisited. Dropping from 2K to 1K on Flash Image saves 33% per image. Every multiplier above 3x in this post comes from a resolution nobody was being billed for, which makes your own request logs the single most valuable thing you can look at this week.

Two mechanical things regardless of which path you pick. Do not follow the pricing page to Gemini 2.5 Flash Image, however tempting $0.039 looks, because that is a second migration in seven weeks. And handle NO_IMAGE before you ship, since the alternative is a silent hole in whatever you generate. You can put your own volumes through our calculator if you want the token arithmetic done against your actual counts rather than our thousand a day.

The documentation trail, contradictions included

All prices come from Google's, OpenAI's, xAI's and Ideogram's own documentation, read on August 10, 2026. Per-image dollar figures for the Gemini models are our arithmetic from published token counts and published per-million rates, and they reproduce Google's own per-image footnotes exactly. No pricing aggregator was used.

  • Google: Imagen models - The shutdown sentence quoted in full, naming August 17, 2026 and Gemini 3.1 Flash Image as the target, plus the three deprecated model IDs
  • Google: Gemini API deprecations - Updated August 3, 2026. The definitions of deprecation and shutdown, the earliest possible dates caveat, the October 2, 2026 date for Gemini 2.5 Flash Image, and the already-dead gemini-3.1-flash-image-preview listed as its replacement
  • Google: Gemini API pricing - Updated August 5, 2026. The $60.00 and $30.00 image output rates, $0.50 input, $3.00 text and thinking output, the per-image footnote for all four resolutions, the 1,290 token figure for 2.5 Flash Image, the $14 per 1,000 grounding rate, the not available free tier rows, and the stale instruction to migrate to Gemini 2.5 Flash Image
  • Google: Vertex AI generative AI pricing - Independent confirmation of the Gemini token rates and the four-resolution per-image footnote. Note that this page no longer carries Imagen rows at all and now presents as Agent Platform pricing, so the Imagen 4 per-image prices in this post were read from the Gemini API pricing page only
  • Google: image generation guide - Thinking enabled by default and not disableable, thinking tokens billed, and up to two interim thought images not billed. Also the source of the Gemini 3 Pro Image token counts, 1,120 at 2K and 2,000 at 4K, which are the figures not to cross-apply to 3.1 Flash Image
  • Google: Batch API - The statement that batch is only available with the generateContent API, which is why Imagen 4 never had a 50% tier and its replacement does
  • Firebase: migrate from Imagen to Gemini models - The shut down as early as August 17 phrasing, and the list of Imagen parameters with no Gemini equivalent including numberOfImages, negativePrompt, imageFormat, personGeneration and addWatermark
  • Google: set output resolution - Updated August 4, 2026. Lists 2K support for imagen-4.0-generate-001 and imagen-4.0-fast-generate-001, so Standard and Fast but not Ultra. This directly contradicts the Gemini API Imagen guide, which says the image-size parameter is supported only for the Standard and Ultra models. The two pages agree only on Standard, which is why the 2K rows in this post are hedged
  • Google AI Developers Forum thread 174305 - Posted July 10, 2026. Independent per-request measurement giving 747, 1,120, 1,680 and 2,520 tokens on generateContent, against a flat 1,120 at every resolution on the Interactions API. A Google representative confirmed the token-count patch on August 3, 2026; the 2K and 4K multi-turn 404 was still open at that reply
  • OpenAI: deprecations - DALL·E removed from the API on May 12, 2026, gpt-image-1 shutting down October 23, 2026, and three further image models dated December 1, 2026
  • Ideogram and xAI API pricing - Ideogram 4.0 Default at $0.060, and grok-imagine-image at $0.020 plus the May 15, 2026 retirement of grok-imagine-image-pro from docs.x.ai. OpenAI's per-image figures come from its own image generation cost table

Three things we could not settle, stated plainly. We do not know how many thinking tokens a typical image request burns, because Google publishes no figure anywhere we could find, so every per-image price here is a floor rather than a total. We could not confirm whether Vertex's Imagen 4 endpoints still answer six weeks after their stated June 30 discontinuation, because checking means a live billed call. And we could not resolve which Imagen 4 tiers actually offered 2K output: two live Google pages give incompatible answers, both updated this year, and no third source arbitrates. We have hedged the 2K rows rather than pick a winner, and flagged the one tier both pages agree on. If you have invoices that settle the thinking-token question, or logs showing 2K coming back from the Fast endpoint, we would like to see them.