Skip to main content
TokenCost logoTokenCost
IndustryJuly 31, 2026·9 min read

OpenAI cut GPT-5.6 Luna by 80% and left Sol untouched. That is a 25x spread inside one model family, and it shrinks to 18x the moment you measure real work.

On July 30 OpenAI took Luna from $1.00 and $6.00 per million tokens down to $0.20 and $1.20, an identical 80% off every line on its card. Terra fell 20% to $2.00 and $12.00. Sol did not move at all, and picked up a Fast mode that charges double. The interesting part is not the discount, it is what the discount did to the distance between OpenAI's own tiers: Sol was 5x Luna on Wednesday and is 25x Luna today. That number is the one everyone will quote, and it is wrong in a way you can measure. Artificial Analysis puts the same evaluation suite through both models and Luna burns 130M output tokens where Sol burns 70M, so the suite costs $3,442.81 on Sol against $190.87 on Luna. Eighteen times, not twenty-five. Roughly 28% of the headline discount is eaten by a model that will not stop talking.

GPT-5.6 Luna pricing after the 80% cut: $0.20 input and $1.20 output per 1M tokens

Photo by Leon Seibert on Unsplash

One announcement, three completely different outcomes

OpenAI opened a GPT-5.6 partner preview on June 26 and took the family generally available on July 9. We wrote up the original three-tier rate card when it first appeared. Three weeks after general availability two of those three tiers are priced differently, and the changes are not proportional to anything. Luna lost four fifths of its price. Terra lost a fifth. Sol lost nothing.

ModelWas (in / out)Now (in / out)Cached inChange
GPT-5.6 Sol$5.00 / $30.00$5.00 / $30.00$0.50No change
GPT-5.6 Terra$2.50 / $15.00$2.00 / $12.00$0.20-20% both sides
GPT-5.6 Luna$1.00 / $6.00$0.20 / $1.20$0.02-80% both sides

The uniformity inside each row is worth pausing on, because it is unusual and it makes your planning easy. Most price changes are lopsided, so working out what happens to your bill means modelling your input-to-output ratio first. Not here. Luna's input, output and cached input all fell by the same 80%, and Terra's all fell by the same 20%. If you were already on Luna, your invoice divides by five, and it does not matter whether you run a prompt-heavy classifier or a generation-heavy chat product. That is a rarer kind of certainty than it sounds.

OpenAI's stated reason is inference efficiency, and the company attributes part of it to GPT-5.6 Sol rewriting production serving kernels and cutting serving cost by around 20%. Take that at face value and the arithmetic still does not reach 80%, which tells you the Luna cut is a positioning decision wearing an engineering story. A 20% efficiency gain buys you Terra's 20%. It does not buy you Luna's.

The floor moved under OpenAI's own floor

Nano is supposed to be the cheapest thing OpenAI sells. It is the tier you drop to when a job is pure classification or extraction and you have decided intelligence is not the binding constraint. GPT-5.4 Nano lists at $0.20 and $1.25. Luna now lists at $0.20 and $1.20, which is to say a current-generation reasoning model with a 1,050,000-token context window is priced fractionally under the previous generation's deliberate floor, and carries more than twice the context of Nano's 400,000.

Whether that is a rounding artefact or a signal, the practical read is the same: there is no longer a cost argument for starting a new project on Nano. The more consequential comparison is external, and the timing is almost comic. Ten days ago Google shipped Gemini 3.5 Flash-Lite at $0.30 and $2.50, which we covered as the third consecutive increase to Google's cheap tier. Luna now undercuts it by a third on input and by more than half on output.

ModelInput / 1MOutput / 1Mvs Luna output
DeepSeek V4-Flash$0.14$0.280.23x
Mistral Small 4$0.15$0.600.50x
GPT-5.6 Luna (new)$0.20$1.201.00x
GPT-5.4 Nano$0.20$1.251.04x
Gemini 3.1 Flash-Lite$0.25$1.501.25x
Gemini 3.5 Flash-Lite$0.30$2.502.08x
GPT-5.4 Mini$0.75$4.503.75x
Claude Haiku 4.5$1.00$5.004.17x

Two things fall out of that ordering. Luna is now the cheapest current-generation reasoning model any large US lab will sell you, which has not been true of an OpenAI model at any point in the last two years. Resist the stronger version of that claim, though, because small non-reasoning models from the same labs still price well underneath it, and the cheapest thing on AWS Bedrock is an order of magnitude below Luna if your work genuinely does not need to think. Luna is also not the cheapest model on the board. DeepSeek V4-Flash charges $0.28 for output against Luna's $1.20, so the gap that actually decides high-volume generation bills is more than four to one and it did not close. If your reason for routing to Chinese models was the output line, July 30 changed nothing for you.

Twenty-five times on paper, eighteen times in practice

Here is the part worth your attention, and it is the reason a rate card is a bad way to choose a model. Artificial Analysis runs one identical evaluation suite against every model it tracks and publishes both the token counts and the dollar total, which means the verbosity of a model and the price of a model land in the same figure. Run that against the GPT-5.6 family and the ladder does not look like the rate card at all.

MeasureSolTerra (max)Luna (max)
Intelligence Index595551
Output tokens on the suite70M96M130M
Cost to run it, before July 30$3,442.81$2,009.91$954.35
Cost to run it, today$3,442.81$1,607.93$190.87
Cost per index point$58.35$29.24$3.74
Output speed65.5 tok/s131.8 tok/s184.4 tok/s
Time to first token147.30s143.74s116.50s

One note on how that table is built, because half of it is arithmetic rather than measurement. The published figures are the current ones. The before-July-30 row is our own calculation, and it is safe to do only because the cuts were perfectly uniform: every Luna line fell by exactly five times and every Terra line by exactly 1.25, so scaling the totals back recovers what the identical suite cost on Wednesday to the cent. Sol never moved, which is why its two rows match. The Terra and Luna runs are the maximum reasoning-effort configurations, which is the expensive end of each model's range.

Now the finding. Across the whole family, only eight index points separate the flagship from the budget tier, 59 down to 51. The rate card asks 25 times more for those eight points. The measured suite asks 18.04 times more, because Luna needs 130M output tokens to do what Sol does in 70M, a 1.86x verbosity penalty that quietly claws back about 28% of the discount OpenAI just announced. Terra is the control case that proves the mechanism. Its card is 2.5x under Sol's and its measured suite comes in 2.14x under Sol's, near enough to agreement, because 96M against 70M is a mild difference in appetite. The further down the ladder you go, the more the sticker lies.

None of which is an argument against the cut. Eighteen times is still an enormous number, and the cost-per-index-point row is the cleanest summary of the day: Luna delivers a point of measured intelligence for $3.74 where Sol asks $58.35, roughly fifteen times better on the metric that blends quality and price. The point is narrower than that. If you budgeted a 25x saving because that is the ratio on the pricing page, you have overbudgeted by about a third, and the gap is entirely explained by tokens nobody quotes.

The last two rows cut the other way and deserve saying out loud, because they invert the usual trade. Luna is not only the cheap tier, it is the fast one: 184.4 tokens per second against Sol's 65.5, and it starts answering half a minute sooner. Buying Sol does not buy you speed. It buys you eight index points, slowly.

The flagship got a more expensive option instead of a cheaper one

Landing the same day, and getting a fraction of the coverage, Fast mode replaces Priority Processing across the GPT-5.6 family. OpenAI describes it as up to 2.5 times faster than standard at twice the price, and it is backward compatible: requests already tagged priority route to Fast mode automatically, so this is a live billing change for anyone using the old flag rather than an opt-in.

ModelStandardBatch / FlexFast modeAbove 272K
Sol$5.00 / $30.00$2.50 / $15.00$10.00 / $60.00$10.00 / $45.00
Terra$2.00 / $12.00$1.00 / $6.00$4.00 / $24.00$4.00 / $18.00
Luna$0.20 / $1.20$0.10 / $0.60$0.40 / $2.40$0.40 / $1.80

Fast mode is a flat 2x multiplier on every tier, which makes it cheaper in relative terms than what it replaced: GPT-5.5 priority ran at $12.50 and $75.00 against a $5.00 and $30.00 list, a 2.5x premium, and we worked through whether that was worth paying in June. The honest framing of July 30 is therefore not that OpenAI cut prices. It cut two tiers, left the third alone, and gave the third a way to charge you double. If you run Sol, the only line that changed for you got more expensive.

The last column is the one that gets forgotten until it appears on an invoice. Cross 272,000 input tokens and the whole request reprices at 2x input and 1.5x output, which is the same cliff GPT-5.5 has always had and it survived the repricing untouched. It is worth noting that even on the wrong side of that cliff, long-context Luna at $0.40 and $1.80 is cheaper than short-context Gemini 3.5 Flash-Lite at $0.30 and $2.50 for any workload where output is more than about a seventh of input.

What it does to an actual monthly bill

Benchmark suites are not production traffic, so here are four shapes we see often, priced at list with no caching, per month. The column that matters is not the saving against old Luna, which is always exactly five times. It is the comparison against whatever you are actually running today.

WorkloadLuna, oldLuna, newGemini 3.5 Flash-LiteHaiku 4.5Sol
Classifier, 200M in / 20M out$320.00$64.00$110.00$300.00$1,600.00
RAG summariser, 100M in / 25M out$250.00$50.00$92.50$225.00$1,250.00
Chat product, 40M in / 40M out$280.00$56.00$112.00$240.00$1,400.00
Agent loop, 500M in / 60M out$860.00$172.00$300.00$800.00$4,300.00

Read those against the verbosity finding rather than on their own. A migration from Haiku 4.5 or Flash-Lite to Luna saves real money on every row here, but the rows assume equal token counts, and Artificial Analysis says Luna is not an equal-token model. Budget the saving at something closer to the 18x pattern than the 25x one and you will not be surprised. Push your own ratios through the cost calculator if your shape looks nothing like these four.

One shape deserves working through rather than tabulating, because the table above hides it. Say you are stuffing a whole repository into context: 400,000 input tokens, 8,000 out, a thousand times a month. That request sits on the wrong side of the 272K line, so it bills at $0.40 and $1.80 rather than $0.20 and $1.20, which is 17 cents instead of 9, and $174.40 a month instead of $87. Painful in isolation. Then price the identical request on Sol, where the same surcharge applies to a much bigger number, and it is $4.36 each, $4,360 a month. The cliff doubles your Luna bill and it doubles your Sol bill, so the ratio between them stays at 25 to 1 and the cheap tier stays cheap. Long context is a reason to watch your prompt sizes. It is not a reason to pick a different tier.

Routing is now a bigger lever than switching providers

Existing Luna traffic needs no attention whatsoever. That bill fell by four fifths on July 30 without a code change, uniformly, and no token mix escapes the benefit. It is the entire population for whom this is unambiguously good news, and the only group that gets to stop reading here.

Teams sitting on Gemini Flash-Lite or Claude Haiku for high-volume, low-judgment work have the first serious reason in months to re-run the comparison. Luna comes in under both on either side of the card, answers faster than the tiers above it, and carries a 1,050,000-token window. Test it against your own traffic before committing, and count output tokens rather than trusting the rate card, because that is precisely where the measured numbers say the surprise lives.

Sol customers got no discount at all, and for them the one line that moved got more expensive. Eight index points now cost $58.35 apiece against Luna's $3.74, and whether that trade is right depends on work only you can evaluate. What has changed is the price of being wrong about it. Choosing between two tiers of the same family is an 18x decision now, which puts it ahead of provider choice in the list of things worth an afternoon. Our pricing table carries the post-cut rates, and the comparison tool will line any two of these up directly.

One admission to close on, because five days ago we argued the opposite. Covering Gemini Flash-Lite we said cheap tokens were getting less cheap and the bottom of the market had stopped falling. Nine days after Google raised its floor, OpenAI went 80% below it. Both observations were accurate on the day they were made, which is the actual lesson: the budget tier has become the fastest-moving part of the price list, in either direction, and no annual plan survives contact with it. Re-check your rates monthly rather than at renewal.

Sources