Skip to main content
TokenCost logoTokenCost
IndustryJuly 3, 2026·8 min read

OpenAI built a chip to make inference cheaper. That is a different thing from making your API bill cheaper.

On June 24 OpenAI and Broadcom pulled the cover off Jalapeño, a custom processor built to run models for less. It joins Google's Ironwood TPUs and AWS Trainium 3 in a race to make a token cheaper to serve. Good news for the labs. The question for you is whether any of it reaches the line item on your invoice, and the honest answer is: only through a route the chip does not control.

Close-up of a processor on a motherboard, representing custom AI inference silicon

Photo by Andrew Dawes on Unsplash

The short version

Five of the biggest names in AI now build their own inference silicon, and Jalapeño is just the newest. The efficiency gains are real: analysts put custom chips at roughly 20 to 50 percent cheaper per useful unit of compute than Nvidia's top parts, and AWS claims five times the tokens per megawatt out of Trainium 3. None of that is the same as a cheaper rate card. A provider's cost to serve a token and the price it charges you are two separate numbers, set by two separate forces. Token prices have fallen hard over three years, but because rivals kept undercutting each other, not because the hardware quietly passed savings along. If you want a smaller bill this quarter, the lever is model choice and competition, not the fab.

What OpenAI actually announced

Jalapeño is OpenAI's first custom chip, co-designed with Broadcom and aimed squarely at inference, the job of running a finished model rather than training a new one. It is a reticle-sized ASIC that OpenAI says went from initial design to tape-out in about nine months, an unusually fast cycle the company credits partly to using its own models in the design loop. First deployment is targeted for the end of 2026, and OpenAI has been testing it on its real-time coding models.

Here is the part worth reading carefully. The official announcement leads with performance per watt being "substantially better than current state-of-the-art," and it talks about low operating cost. It does not put a dollar-per-token number on the table, and it does not promise that your API price drops by any amount. The widely repeated "50 percent cheaper per token" line traces to Broadcom's CEO citing early lab testing, not to OpenAI's written release, and it comes with no disclosed baseline or independent benchmark. That gap between what a chip does in a data center and what shows up on your bill is the whole story here.

Everyone is building inference silicon now

OpenAI is late to a party the other hyperscalers have been throwing for a while. Each of these chips is pitched first as an inference part, and each headline efficiency figure is the maker's own claim, measured against its own baseline rather than a shared benchmark. Read the column as direction and intent, not a leaderboard.

ChipMakerHeadline efficiency claimStatus
JalapeñoOpenAI + BroadcomPerf/watt "substantially better than state of the art" (no cost figure given)Deploying late 2026
Ironwood (TPU v7)Google2x perf/watt vs prior TPU; built "for the age of inference"Generally available
Trainium 3AWS4x perf/watt vs Trainium 2; ~5x output tokens per megawattGenerally available
Maia 200Microsoft30% better performance per dollar vs prior fleet hardwareDeploying
MTIA (v3+)Meta"More cost efficient" than general-purpose chips (no number given)In production

The pattern is the same everywhere: custom parts built to serve tokens at a lower cost per watt than a general-purpose Nvidia GPU. Independent analysts at SemiAnalysis reckon Google's TPUs land somewhere around 20 to 50 percent lower cost per useful FLOP than Nvidia's GB200 and GB300 class, and Anthropic now runs Claude across a mix of Trainium, TPU, and GPU rather than betting on one. So the supply side is genuinely getting cheaper to operate. Hold that thought against your last invoice.

Token prices really have fallen off a cliff

Before arguing that hardware savings do not reach you, it is only fair to admit that prices have dropped, and dropped hard. For a fixed level of capability, both a16z and Sam Altman independently landed on the same rule of thumb: about 10x cheaper per year. Epoch AI, measuring the cost to hit a fixed benchmark score, found declines running anywhere from 9x to 900x per year depending on the task, with a median near 50x. That is faster than anything Moore's law ever delivered.

The flagship sticker prices tell a slower, more honest version of the same story, and they hide a detail worth seeing.

FlagshipLaunchInput / 1MOutput / 1M
gpt-4Mar 2023$30.00$60.00
gpt-4oMay 2024$2.50$10.00
gpt-5Aug 2025$1.25$10.00
gpt-5.6-lunaJun 2026 (preview)$1.00$6.00

Input fell about 30x in three years. Output fell 6x, then sat dead still at ten dollars from GPT-4o all the way through GPT-5, more than a year with no movement on the number that dominates most bills. Only the GPT-5.6 Luna preview finally nudged output below it. If cheaper silicon translated directly into cheaper output, that flat stretch would not exist. Something else sets the price. For the longer arc of this collapse, we charted it in the AI price index.

Why the chip in the rack is not the price on your invoice

A provider's cost to serve a token sets the floor under its price. It does not set the price. Three things sit in between, and each one absorbs the hardware savings before they can reach you.

First, labs price to value, not to cost. A model that can close a support ticket or ship a code change is worth far more than the pennies of compute behind it, so a cheaper chip mostly widens the gap between cost and price. That gap has a name, margin. Independent estimates put Anthropic's gross margin on Claude served over Trainium north of 60 percent, which is exactly what you would expect when efficient in-house silicon meets value-based pricing. The efficiency is real and it is landing in the margin, not the rate card.

Second, the models got hungrier. A reasoning or agentic model can spend tens of thousands of internal tokens chewing on a single request before it answers. Even if the per-token rate drops, a task that used to cost you 800 tokens and now costs 40,000 comes out more expensive. We walked through exactly this trap with GPT-5.5 token efficiency: the sticker looked flat while real per-task cost climbed.

Third, demand keeps eating the slack. Every efficiency gain frees capacity, and that capacity gets poured straight into bigger, more capable models rather than lower prices on the current ones. When a lab spends tens of billions on TPUs and Trainium clusters, it is not planning to serve today's model for less. It is planning to serve a much larger one at a similar price. Cheaper inference funds the next frontier model; it does not refund you.

What actually cuts your bill

Prices do fall, so something works. It is competition, and it arrives in shocks rather than a smooth glide. When DeepSeek R1 shipped in January 2025 at roughly 27x under OpenAI's comparable reasoning model, it did more to reset the market in a week than a year of quiet efficiency gains had. The incumbents cut because a rival made them, and we tracked one such round in the OpenAI versus Anthropic price war. Cheaper chips are the ammunition that makes those cuts survivable for the lab. They are not the trigger.

The lever you actually hold is model choice. Moving a workload off a flagship and onto a competent mid-tier or open-weight model routinely saves more than any single vendor price drop, and it takes effect the moment you change a string in your config. Trimming output tokens, caching repeated context, and batching non-urgent calls stack on top of that. Waiting for Jalapeño to lower your bill is a bet on someone else's pricing committee. Repricing your own stack is a bet you control. Run your real token mix through the cost calculator and the cheaper path usually shows up fast.

Compare every model pricePrice your current stack

FAQ

What is OpenAI's Jalapeño chip?

OpenAI's first custom inference processor, co-designed with Broadcom and unveiled June 24, 2026. It is a reticle-sized ASIC built to run models, taken from design to tape-out in about nine months, with first deployment targeted for late 2026. OpenAI leads with performance per watt, not a dollar-per-token cut.

Will custom AI chips make LLM API prices cheaper?

Eventually and indirectly, not automatically. Cheaper silicon lowers a provider's cost to serve a token, which mostly widens margin. Your price is set by competition and value, so it drops when a rival forces it, as DeepSeek did in early 2025, not because the hardware got more efficient.

How much have LLM API prices actually fallen?

For a fixed capability, roughly 10x per year per a16z and Sam Altman, and 9x to 900x per year to hit a fixed benchmark per Epoch AI. Sticker prices moved slower: GPT-4 at $30/$60 in 2023, GPT-4o at $2.50/$10 in 2024, GPT-5 at $1.25/$10 in 2025. Input collapsed while output stayed sticky.

Why does cheaper hardware not lower my bill one-for-one?

Labs price to the value of the output, so savings become margin. Reasoning models burn far more tokens per task, so per-request cost can rise even as the per-token rate falls. And demand keeps outrunning supply, so freed capacity funds bigger models instead of lower prices.

What actually lowers my API costs then?

Competition and your own model choice. A cheaper rival forces incumbents to cut, and moving a workload from a flagship to a mid-tier or open-weight model usually saves more than any single price drop. Caching, batching, and trimming output tokens compound on top.

Sources