grok-voice-latest moved to Think Fast 2.0 this morning, so unchanged code now bills 60% more per minute. The sentence before the price reads "we believe pricing should be predictable and transparent."
The alias repointed today. Call grok-voice-latest this morning and you are on Think Fast 2.0 at $0.08 a minute of audio instead of 1.0 at $0.05, which turns $3.00 an hour into $4.80 for a codebase that did not change a character. xAI published the date and published the new price, so none of this is hidden. It just never published the old price next to the new one, never used the word increase, and has no deprecation policy page to put this under. What the money buys is real: 75.7% to 82.9% on Artificial Analysis' Speech to Speech Index, and time to first audio nearly halved to 0.70s. Two things I did not expect, though. That 82.9% is second place, behind an Alibaba model that scores higher and costs less per hour. And the single number you would need to budget any of this, whether a two-way minute meters once or twice, is not documented anywhere. Resolve it one way and Grok is 17% cheaper than OpenAI's flagship. Resolve it the other and it is 67% dearer.

Photo by Ivan Jermakov on Unsplash
Three dollars an hour became four eighty
Both prices and the migration date are first-party, off xAI's pricing page and its release notes, so there is no inference in the headline number. The benchmarks below are Artificial Analysis', but read off AA's own page rather than out of xAI's announcement table, which turns out to matter. Later on I convert OpenAI and Google into per-minute figures using each vendor's own documented audio density, never one vendor's ratio applied to another.
xAI runs several audio SKUs and exactly one of them moved. Worth seeing them together, because the conversational line now sits far above the transcription lines, and plenty of voice products are paying speech-to-speech rates for work a $0.10-an-hour endpoint would do.
| xAI audio SKU | Billing unit | Rate | Per hour |
|---|---|---|---|
| grok-voice-think-fast-2.0 | Minute of audio sent or received | $0.08 | $4.80 |
| grok-voice-think-fast-1.0 | Minute of audio sent or received | $0.05 | $3.00 |
| Text input, either version | Per conversation.item.create event | $0.004 | n/a |
| Speech to text, streaming | Hour | $0.20 | $0.20 |
| Speech to text, REST | Hour | $0.10 | $0.10 |
| Text to speech | 1M characters | $15.00 | n/a |
That text line deserves a second look, because it is not a token price and reading it as one will mislead you in both directions. xAI charges $0.004 for every conversation.item.create event the client sends, flat, whatever the length. A five-token nudge costs the same as a five-thousand-token context injection. Tool results come back free, audio items bill on the audio meter, and response.create is not billable at all. So a chatty orchestration layer that pokes the session 250 times has spent a dollar before anyone speaks, while a single enormous system prompt is four tenths of a cent. Nothing else on this site prices that way.
Here is what the audio change does to a bill, assuming the charitable metering reading. Nothing about this increase responds to prompt discipline or caching or any lever people normally reach for. It is a flat 1.6x on audio, applied on a date, to whoever typed an alias instead of a version.
| Monthly audio on the alias | Billed yesterday | Billed today | Difference |
|---|---|---|---|
| 1,000 min, a pilot | $50 | $80 | +$30 |
| 10,000 min | $500 | $800 | +$300 |
| 60,000 min, 1,000 hours | $3,000 | $4,800 | +$1,800 |
| 250,000 min, a real call centre | $12,500 | $20,000 | +$7,500 |
One small thing that says a lot about how this shipped. On migration morning, xAI's canonical speech-to-speech model page still reads "starting at $0.05 / minute or $3.00 / hour" and carries a last-updated date of July 27, two days before 2.0 was announced. The $0.08 exists on the pricing page and in the blog post. The page a developer lands on from the API reference has not caught up.
Second place, and the model above it is cheaper
The quality jump is real and I want to be fair about it before being unkind about the framing. Artificial Analysis benchmarked 2.0 on release day and it moved 7.2 points, which is a lot for a point release.
| Measure | 1.0 | 2.0 | Best on the board |
|---|---|---|---|
| Speech to Speech Index | 75.7% | 82.9% | Qwen Audio 3.0 Plus, 84.1% |
| Big Bench Audio, reasoning | 97.1% | 97.2% | Qwen Audio 3.0 Plus, 99.2% |
| Full Duplex Bench, dynamics | 77.8% | 95.1% | Qwen Audio 3.0 Plus, 98.4% |
| ๐-Voice, agentic | 52.1% | 56.5%, first | Grok Voice 2.0 |
| Time to first audio | 1.25s | 0.70s | Deepslate Opal, 0.44s |
| Price per audio minute | $0.05 | $0.08 | Gemini 3.1 Flash, $0.025 |
Now the framing. xAI's announcement compares 2.0 against OpenAI and Google and stops there, which produces a clean sweep. Put Alibaba back in and 2.0 is second on the index, second on reasoning, second on conversational dynamics, and first on the one component that measures finishing a customer service task. Qwen Audio 3.0 Realtime Plus takes 84.1% at $4.42 an hour of input audio against Grok's $4.80: higher score, 8.5% cheaper. That comparison is absent from every write-up I read, because every write-up worked from the vendor's table.
The honest case for Grok is latency, and it is a strong one. 0.70s to first audio against Qwen's 4.02s is a 5.7x gap, and it is the only model in the top five that averages under a second. In a phone call that difference is the whole product. Nobody abandons a call over 1.1 index points; they abandon it over four seconds of silence. Which makes it odd that latency is the axis xAI buried and the index is the one it led with.
Divide price by quality and the upgrade looks worse than it reads. Price rose 60%, index quality rose 9.5%, so cost per index point went from $0.0397 to $0.0579 per point-hour, up 46%. That framing is a little unfair, since points get harder to buy as you climb and latency does not appear in the index at all. It is still the direction of travel: 1.0 was the value pick in this category and 2.0 is not.
The efficiency gain is denominated in a unit you are not billed in
xAI makes a point of saying 2.0 needs roughly 0.4x the reasoning tokens of 1.0 for the same work. On a token-priced model that would be the headline and it would offset most of a 60% rate rise by itself.
This model is billed per minute of audio. Reasoning tokens never reach the invoice. What fewer of them buys is the 0.70s time to first audio, and the only path from there to a smaller bill is indirect: less dead air per turn, marginally shorter calls, a few fewer metered minutes. I cannot put a number on that and neither can xAI, because it depends on whether your callers fill the silence or wait through it. A support line where the human keeps talking sees none of it.
I keep running into this shape and it is worth naming. Vendors report efficiency in tokens, because tokens are what they optimise, then bill in a unit tokens do not map onto. It happened in reverse on Qwen3.8-Max, where output priced 60% under Kimi K3 produced a bill only 11% smaller because the model wrote 2.4x as many tokens. Same lesson, opposite direction: the unit on the rate card is rarely the unit your workload varies in.
Minutes and tokens are not the same product
To judge whether $0.08 is a lot you have to get OpenAI and Google onto the same axis, and they do not sell audio by the minute. They sell it by the token, at a documented density, and the two of them use neither the same number nor the same shape.
OpenAI is asymmetric. One audio token per 100ms of user speech, one per 50ms of model speech, so 600 tokens for a minute of you and 1,200 for a minute of the model. Output audio also costs twice as much per token, which means model speech runs 4x the price of your speech second for second. Google is symmetric: 25 tokens per second in both directions on the Live API. That asymmetry is why the two cost metrics in the next section disagree.
Before the table: the one number nobody publishes
xAI defines the audio meter as covering "audio sent or received." One meter, both directions. It does not say whether a minute of conversation with 30 seconds of speech each way bills as one minute of audio or two, and there is no worked billing example anywhere in the docs. I checked the pricing page, the model page, the speech-to-speech guide and the Realtime API reference.
That unanswered question is worth 2x on every Grok row below, which is more than the distance between any two vendors in the table. So the table carries both readings and picks neither. If you are budgeting against this model, it is the first thing to ask xAI and the last thing to assume.
| Option | Your 60s | Its 60s | A minute each way | Per hour |
|---|---|---|---|---|
| grok-voice-2.0, if both streams meter | one meter | one meter | $0.1600 | $9.60 |
| gpt-realtime-2.1 | $0.0192 | $0.0768 | $0.0960 | $5.76 |
| grok-voice-2.0, if wall clock | one meter | one meter | $0.0800 | $4.80 |
| ElevenLabs Agents | call duration | call duration | $0.0800 | $4.80 |
| Deepgram Voice Agent | call duration | call duration | $0.0750 | $4.50 |
| Cartesia Line | call duration | call duration | $0.0600 | $3.60 |
| grok-voice-1.0, if wall clock | one meter | one meter | $0.0500 | $3.00 |
| gpt-realtime-2.1-mini | $0.0060 | $0.0240 | $0.0300 | $1.80 |
| gemini-3.1-flash-live | $0.0045 | $0.0180 | $0.0225 | $1.35 |
Read the top and third rows together, because they are the same product under two readings of one sentence. If xAI meters wall-clock call duration, Grok Voice 2.0 at $0.08 undercuts OpenAI's flagship by 17% even after today's rise. If it meters both speech streams, the same conversation costs $0.16 and Grok is 67% more expensive than OpenAI. Every other per-minute vendor in that table says "call duration" or "minute of conversation" in its own docs. xAI says "audio sent or received," which is the one phrasing that leaves the door open, and it publishes no worked example to close it.
The other thing worth staring at is the bottom row. Google's native audio does the same job for $0.0225 a minute, which is 3.6x under Grok on the charitable reading and 7.1x under it on the other. Gemini 3.1 Flash Live scores 69.5% against Grok's 82.9%, so it is genuinely a lesser model, and its 2.99s time to first audio is four times slower. For an agent holding a plan in its head, pay the difference. For a deflection bot reading from a script, that gap is very hard to justify.
Two caveats before anyone quotes this table at a vendor. The ElevenLabs figure excludes the LLM and the telephony, both billed separately, so it is not all-in the way xAI's single meter is. And every row assumes a conversation split evenly between the speakers. Move that ratio and the token-priced rows move while the flat-rate rows do not.
Two published metrics, two opposite verdicts
Artificial Analysis publishes a cost column beside its index, and it does not agree with my table. Worth walking through, because the disagreement is instructive rather than an error in either.
| Rank | Model | Index | $/hr input audio | TTFA |
|---|---|---|---|---|
| 1 | Qwen Audio 3.0 Realtime Plus | 84.1% | $4.42 | 4.02s |
| 2 | Grok Voice Think Fast 2.0 (High) | 82.9% | $4.80 | 0.70s |
| 3 | GPT-Realtime-2.1 (High) | 79.1% | $10.75 | 1.22s |
| 4 | GPT-Realtime-2 (High) | 77.2% | $4.14 | 1.14s |
| 5 | Qwen Audio 3.0 Realtime Flash | 76.3% | $4.77 | 4.16s |
| 6 | Grok Voice Think Fast 1.0 | 75.7% | $3.00 | 1.25s |
| 7 | Gemini 3.1 Flash (High) | 69.5% | $1.75 | 2.99s |
| 11 | Deepslate Opal | 62.8% | $6.48 | 0.44s |
By this column Grok Voice 2.0 is the second most expensive thing in the top five and dearer than the model that beats it. But look at rows three and four: GPT-Realtime-2.1 and GPT-Realtime-2 have identical rate cards, $32 audio input and $64 audio output per million tokens, and AA prices one at $10.75 an hour and the other at $4.14. Nothing about the price changed. What changed is how much the model says, because AA measures the cost of completing a fixed benchmark and normalises it by the length of the input audio. A chattier model bills more against the same denominator.
That is the whole disagreement. AA prices per hour of input audio, and OpenAI's expensive leg is output. For Grok Voice the distinction collapses, because $3.00 an hour was simply $0.05 a minute with the units changed, and AA's $4.80 for 2.0 is $0.08 times sixty. For OpenAI the distinction decides the answer. Neither number is wrong and neither is a bill.
So the answer to "is Grok Voice expensive now" depends on who does the talking in your product. Build something that mostly listens and OpenAI looks better than its sticker. Build something chatty and the flat minute gets more attractive the more the model says. Nobody publishes that ratio for your application, which is the argument for measuring it before you migrate anything. The underlying rates are on our model comparison pages and you can run your own volumes through the cost calculator.
What a flat minute buys that a token does not
There is a structural difference under the sticker prices that almost nobody puts in a comparison table, and having read the docs properly I think it matters more than the stickers do.
OpenAI's realtime cost guide says the entire conversation is sent to the model on every response, and that turns later in a session therefore cost more than turns earlier in it. So the $0.096 in my table is a floor for the first minute and nothing else. Minute nine of the same call costs more than minute one, because minute one is still in the prompt. Google hits the same accumulation and solves it with hard stops rather than a bill: the Live API caps sessions at 15 minutes for audio only, or two minutes with video, unless you turn on context window compression.
A flat per-minute meter has no such drift. Minute nine costs what minute one costs, and a ten-minute call is exactly ten times a one-minute call. Anyone who has tried to forecast a realtime bill from a token rate knows what that predictability is worth, and it is the real argument for xAI, ElevenLabs, Deepgram and Cartesia over the token-priced APIs. It also means comparing $0.08 against $0.096 understates xAI's case for long calls and overstates it for very short ones, where OpenAI's cached audio input at $0.40 against $32 does the most work.
xAI caps sessions too, at 120 minutes, with ten concurrent sessions per team and us-east-1 as the only region. Voice rate limits sit outside the published tiers and go through sales. And the context window for either voice model is not documented anywhere I could find, which is a strange gap for a model sold on holding a conversation.
An alias is a standing instruction to accept the next price
The immediate job takes one line. If your voice traffic goes through grok-voice-latest and you have not decided you want 2.0, change it to grok-voice-think-fast-1.0 and the rate goes back to $0.05. Then choose on the merits rather than on a default somebody else picked.
Worth knowing how the docs guide you here, because they pull in two directions on the same page. The model table says the alias points at 1.0 while the prose two sections down says it currently points at 2.0. One paragraph tells you to use the alias for new integrations so your app tracks the recommended model; the next tells you to pin a version in production for stability. The Realtime API reference lists the alias as the default and recommends it. And xAI's general model page says aliases suit users who want the latest features, without mentioning that features can arrive attached to a higher rate. There is no deprecation policy page and no price-change notification policy anywhere on the docs site. I looked for both.
On whether to keep 2.0, check one number in your own logs before any benchmark: how much of each call is the model speaking. If the answer is most of it, 2.0 at a flat rate is a reasonable buy and the latency improvement is real money in a live conversation. If your product mostly listens, you are paying conversational rates for transcription, and either Gemini 3.1 Flash Live at $0.0225 a minute or xAI's own $0.10-an-hour speech-to-text endpoint does that job for a fraction. If you want the top of the index and can live with four seconds of silence, Qwen Audio 3.0 Realtime Plus scores higher and costs less.
The thing I would not do is trust any single per-minute figure, mine included, without knowing which direction of audio it counts. That ambiguity is doing a lot of quiet work across this whole category, and it is why two careful sources can look at the same model on the same morning and disagree about whether it is the cheap option or the dear one.
Sources
- xAI: pricing - The primary source for grok-voice-think-fast-1.0 at $0.05/min ($3.00/hr) and 2.0 at $0.08/min ($4.80/hr), the $0.004 text input line on both, speech-to-text at $0.10/hr REST and $0.20/hr streaming, and text-to-speech at $15.00 per 1M characters. Note the page renders the text fee as bare "$0.004 / text input" with no denominator
- xAI: speech-to-speech model page - Where the audio meter is defined as covering "audio sent or received", the $0.004 is defined as a flat fee per conversation.item.create event with function_call_output and audio items excluded, and response.create is confirmed non-billable. Also the 120-minute session cap, 10 concurrent sessions per team and us-east-1 region. Still reads "starting at $0.05 / minute" and dated July 27, two days before 2.0 launched
- xAI: speech-to-speech guide - Carries the alias table saying grok-voice-latest updates to 2.0 on August 5, 2026, and the contradicting prose saying it "always points to the newest model (currently grok-voice-think-fast-2.0)", plus the advice to use the alias for new integrations and to pin in production. Source for the 20 documented languages and the note that bare "es" and "pt" are rejected
- xAI: developer release notes - The July 29 entry announcing 2.0 and the August 5 alias routing, with no price attached. For contrast the July 8 Grok 4.5 entry on the same page does quote $2 and $6 per million, so omitting the voice price was a choice rather than a format limit
- xAI: introducing Grok Voice Think Fast 2.0 - The July 29 announcement. Source for 82.9% on the Speech to Speech Index, Big Bench Audio 97.2%, Full Duplex Bench 95.1%, ๐-Voice 56.5%, 0.70s time to first audio, the 0.4x reasoning-token claim, the instruction to pin 1.0 before August 5, and the "no action needed to upgrade" wording. Its comparison table includes only OpenAI and Google. Returns 403 to a direct fetch and was read through a text-extraction proxy
- Artificial Analysis: speech-to-speech leaderboard - The full ranking and cost column this post uses, read out of the page's embedded datasets because the tables render client-side: Qwen Audio 3.0 Realtime Plus 84.07% at $4.42/hr and 4.019s, Grok Voice 2.0 82.93% at $4.80 and 0.70s, GPT-Realtime-2.1 High 79.06% at $10.75 and 1.215s, GPT-Realtime-2 High 77.22% at $4.14 and 1.14s, Qwen Audio 3.0 Realtime Flash 76.33% at $4.77, Grok Voice 1.0 75.67% at $3.00 and 1.247s, Gemini 3.1 Flash High 69.54% at $1.75, Deepslate Opal 62.76% at $6.48 and 0.439s. Also the ๐-Voice ordering that puts Grok 2.0 first at 56.48% and Qwen Plus second at 54.60%
- Artificial Analysis: Speech to Speech Index methodology - The index is an equal third each of Big Bench Audio, Full Duplex Bench and ๐-Voice, and the cost column is defined as the cost of completing a fixed 40-question Big Bench Audio subset normalised by the length of input audio. That definition is why two models on an identical rate card can differ 2.6x in the cost column
- OpenAI: API pricing - gpt-realtime-2.1 at $32.00 audio input, $0.40 cached audio input and $64.00 audio output per 1M tokens, with gpt-realtime-2.1-mini at $10.00 and $20.00. Also the OpenAI voice SKUs that are billed per minute rather than per token, gpt-realtime-translate at $0.034/min and gpt-live-transcribe at $0.017/min, which is how you know minute-billing is a product decision rather than a house style
- OpenAI: realtime cost guide - The audio density this post converts with, one token per 100ms of user audio and one per 50ms of assistant audio, plus the statements that the entire conversation is resent on every response so later turns cost more, and that cache hits are best-effort and bust on any edit to history
- Google: Gemini API pricing - gemini-3.1-flash-live-preview at $3.00/1M or $0.005/min audio input and $12.00/1M or $0.018/min audio output, and the note that Live API billing is calculated at 25 tokens per second of audio. Publishing both units is what let us verify the conversion instead of assuming it
- Google: Live API best practices - Restates 25 tokens per second and documents the session ceilings that context accumulation forces, 15 minutes audio-only and two minutes with video unless context window compression is enabled
- ElevenLabs: Agents pricing - $0.08/min overage at every tier, $0.16/min past the concurrency limit, and the caveat that the LLM and telephony bill separately, so it is not all-in comparable to xAI's single meter
- Deepgram: pricing - Voice Agent API at $0.075/min pay-as-you-go standard, $0.065 bringing your own text-to-speech, and $0.163 on the advanced tier
- Cartesia: pricing - Line at $0.06 per minute of call duration, plus $0.014/min if the number is Cartesia's. Quoted here because the phrase "call duration" is exactly the clarity xAI's meter is missing