Skip to main content
TokenCost logoTokenCost
Model ReleaseAugust 11, 2026·11 min read

Meta gave Muse Glimmer's weights away yesterday and exactly one company sells it. $0.35 and $1.50 is not a market price, it is Together AI's price, and the only thing testing it is a 15.9 GB file you can download instead.

Meta released a 30B dense model under Apache 2.0 on August 10, its first open weights in more than a year, and named twelve launch partners. Nine of them are local runtimes. Of the three that sell hosted inference, Fireworks never shipped a listing, and OpenRouter turns out to be a storefront: pull its endpoints API for the model and you get a single entry, provider name Together, at the same two numbers Together prints on its own site. So the entire published market for Muse Glimmer inference is one rate card. That matters more than it sounds, because open weights usually do the opposite. gpt-oss-120b sells from $0.03 to $0.35 per million input across its hosts and Gemma 4 31B from $0.08 to $0.99, spreads of 11.7x and 12.4x that exist only because hosts undercut each other. Muse Glimmer's spread is 1.0x. Price a month of ordinary agent traffic and it comes to exactly $20.00, against $1.88 for gpt-oss-120b and $4.60 for Gemma 4 31B, both also Apache 2.0, and $12.80 for GPT-5.6 Luna, which is not open at all. The competitive floor here is not another host. It is a 24 GB graphics card.

Abstract network of blue lines and glowing nodes from Meta's Muse Glimmer announcement

Image source: Meta AI Research

What you are actually buying

A capable small model, priced by a monopoly of one, with a free escape hatch sitting next to it. If you rent it, you are paying more than several open-weight peers charge for comparable work and you have nowhere to fail over to. If you can run a 24 GB card, the licence means the rental price is a ceiling rather than a rate. Neither of those is a reason to avoid the model. Both are reasons to not put $0.35 into a spreadsheet and stop thinking.

Hosts with a public price

1

Together AI. The rest resell it

Price spread across hosts

1.0x

gpt-oss-120b runs 11.7x

Runs on

15.9 GB

4-bit, fits a 24 GB card

Hallucination rate

82%

Qwen3.6 27B measures 49%

The launch partner list has one seller in it

Meta named twelve launch partners, and the shape of that list is the story. Nine are local runtimes and serving frameworks: Ollama, LM Studio, Unsloth, llama.cpp, ExecuTorch, MLX, vLLM, SGLang and PyTorch. Only three sell hosted inference, and a day later that trio resolves to a single company taking money for tokens.

WhereInput / 1MOutput / 1MWhat it really is
Together AI$0.35$1.50The host. Everything below traces here
OpenRouter$0.35$1.50One endpoint, forwarded to Together
ModelsLab$0.35$1.50Reseller, identical card
NVIDIA build.nvidia.comNo priceNo priceFree trial tier, no published per-token rate
Fireworks AINot listedNot listedNamed as a partner, no listing yet
MetaNot sellingNot sellingPoints you at Hugging Face instead

OpenRouter is usually worth routing through precisely because it spreads a model across hosts and fails over when one of them wobbles. Here it has nothing to fail over to. Its own page says the model is hosted by one provider and that it forwards every request directly. A Together outage and an OpenRouter outage are the same outage.

The absence worth noting is Meta's. It shipped a model with an agentic pitch and no endpoint of its own, not even a rate-limited free tier, which is a different posture from Muse Spark 1.2, where Meta sells the tokens itself and discounts them heavily in exchange for training rights. Meta has said nothing either way about a hosted Glimmer API. Treat that as unaddressed rather than ruled out. It was the first thing the Hacker News thread on the launch picked at too, where the top comment reads as surprise that Meta will not host this even as a rate-limited free tier.

What Meta actually shipped

SpecValue
Parameters~29.6B dense, incl. a ~1.8B vision encoder. Not MoE
LineageLogit-distilled from Muse Spark during pre-training
LicenceApache 2.0, not a Llama community licence
Context131,072 tokens
Max outputNot published by anyone
ModalitiesText and image in, text out. 100+ languages
Knowledge cutoffJanuary 4, 2026

Two numbers in circulation are wrong and worth correcting before they end up in someone's capacity plan. Artificial Analysis' model page lists a 256K context; the config, Together's endpoint and NVIDIA's NIM docs all say 131,072. And a 32,768 figure floating around the Hugging Face blog is a value inside an example agent config, not a model limit. On max output tokens, nobody publishes one: Together returns null for it. If you are sizing a job around a hard output cap, measure it rather than quoting it.

The licence is the part that changes the economics. Apache 2.0 is not the Llama community licence Meta shipped Scout and Maverick under, with its user-count threshold and naming conditions. There is no clause to read here, which is why the self-hosting arithmetic further down is a genuine alternative rather than a thought experiment.

The same month of work: $20.00 here, $1.88 on another Apache 2.0 model

One workload, priced everywhere. Take a modest agent that consumes 40M input tokens and emits 4M output tokens over a month, roughly a 10:1 ratio, which is what tool-heavy loops tend to look like. Open models are priced at the cheapest single host that sells both legs, since that is what you would actually pay. That distinction matters: take the lowest input from one host and the lowest output from another and you get a number nobody can buy.

ModelWeightsIn / Out per 1MMonthly bill
Qwen3.7 FlashClosed$0.03 / $0.13 ≤32K$1.72
gpt-oss-120bApache 2.0$0.03 / $0.17$1.88
Mistral Small 3.2 24BApache 2.0$0.075 / $0.20$3.80
Gemma 4 31BApache 2.0$0.08 / $0.35$4.60
Llama 4 MaverickLlama 4 community$0.20 / $0.696$10.78
GPT-5.6 LunaClosed$0.20 / $1.20$12.80
Qwen3.6 27BApache 2.0$0.30 / $2.00$20.00
Muse Glimmer 30BApache 2.0$0.35 / $1.50$20.00
Gemini 3.5 Flash-LiteClosed$0.30 / $2.50$22.00
Claude Haiku 4.5Closed$1.00 / $5.00$60.00

Muse Glimmer lands eighth of ten, level to the cent with Qwen3.6 27B. Run the same traffic through gpt-oss-120b and you pay $1.88, which is nine cents on the dollar, and gpt-oss-120b is also Apache 2.0, also downloadable, and considerably larger. Gemma 4 31B, its closest peer by size and licence, does the job for $4.60. And GPT-5.6 Luna comes in $7.20 under it, which is a proprietary model from a lab with no incentive to be generous undercutting an open-weights release by more than a third.

Read the spread column rather than the price column and the reason becomes obvious.

Open modelInput range across hostsSpread
Gemma 4 31B$0.08 to $0.9912.4x
gpt-oss-120b$0.03 to $0.3511.7x
Qwen3.6 27B$0.289 to $0.602.1x
Llama 4 Maverick$0.20 to $0.351.75x
Mistral Small 3.2 24B$0.075 to $0.0941.25x
Muse Glimmer 30B$0.35 to $0.351.0x

Every other row on that table is a number that got argued down. Somebody listed Gemma 4 31B at $0.99, somebody else listed it at $0.08, and the second host exists because the first one left room. Note where Muse Glimmer's single price lands: $0.35 is exactly the top of gpt-oss-120b's range and 4.4x the bottom of Gemma 4's. It is not priced like an outlier. It is priced like a reasonable opening bid that nobody has answered yet.

This is the opposite of what we found on Ling-3.0-flash, where four hosts produced three different cards and a 3.5x swing on identical work. There the problem was picking the right host. Here there is no picking to do.

15.9 GB is the ceiling on that price

A 30B dense model quantises down small enough to be a consumer problem rather than a data-centre one. The community GGUF conversions are the only file sizes anyone can observe directly, so those are what we are using.

QuantisationSizeCard it fits
4-bit (UD-Q4_K_XL)15.9 GB24 GB RTX 4090, comfortably a 32 GB 5090
5-bit19.2 to 21.8 GB32 GB for real headroom
6-bit26.3 GB32 GB
8-bit (Q8_0)29.6 GB40 GB, or 48 GB with context
BF1655.7 GB1x H100 80GB

Sources disagree on the full-precision figure. Meta says 55 GB, the actual BF16 shards total 55.7 GB, Artificial Analysis says about 60, and one write-up says 59.55. Take the range as 56 to 60 and stop worrying about it, because nobody renting a single H100 cares which end they land on. The 4-bit number is the one that matters and it is observable: 15.9 GB, plus roughly 1.8 GB of KV cache at minimum.

Now the arithmetic that decides whether renting is sensible. At $1.50 per million output tokens, a graphics card that costs $2,500 has paid for itself once it has produced 1.67 billion output tokens. An RTX 5090 is reported at 74.9 tokens per second on this model, and 233.4 with DFlash speculative decoding, a 3.1x improvement. At the fast figure, 1.67 billion tokens is about 83 days of continuous generation. At the baseline figure it is about 258 days. Electricity over that fast run adds roughly 1,140 kWh, about $171 at $0.15 per kWh, so it moves the answer by a week rather than changing it.

The $2,500 is our assumption, not a quoted price, and those decode figures are single-stream. Batch the card properly and the payback gets shorter, not longer, so the conclusion holds in the direction that matters: if you are generating output around the clock, the hardware wins inside three months. If you are running a few million tokens a month, $20.00 is cheap and you should not be shopping for a GPU. What you should not do is assume the rented price is a market price, because there is no market. Our Kimi K3 self-hosting breakdown has the GPU-hour method in full if you want to run your own numbers.

Meta picked the benchmarks and won five of the nine

The model card puts Muse Glimmer against Qwen3.6-27B and Gemma4-31B. Read it row by row rather than taking the summary, and the margins are thinner than the framing suggests.

BenchmarkMuse GlimmerQwen3.6-27BMargin
MCP Atlas75.562.5+13.0
DeepSearch QA74.671.1+3.5
SWE-Bench Pro51.250.2+1.0
AIME 202694.794.1+0.6
CharXiv Reasoning78.878.4+0.4
MMMU Pro7475-1.0
SWE-Bench Verified76.077.2-1.2
TerminalBench 2.151.760.7-9.0
OSWorld-Verified65.975.6-9.7

Five wins and four losses, on Meta's own selection. One of the wins is real and it is the one Meta is pitching: 13 points on MCP Atlas, which is agentic tool use. After that it is 3.5 on DeepSearch QA and then three margins of a point or less. The losses run the other way, two of them close and two of them around nine points, on terminal work and computer use. SiliconANGLE reports Meta ran roughly two dozen benchmarks and took first place in about half, which is what a nine-row table selected from two dozen tends to look like.

Worth flagging one number that is circulating wrongly: several write-ups have Muse Glimmer at 60.7 on TerminalBench 2.1. That is Qwen's score. Meta's card says 51.7, and Artificial Analysis independently measures 52%, so the two agree and the higher figure is a transcription error rather than a harness disagreement.

Artificial Analysis, running its own suite, reaches the opposite overall verdict. Intelligence Index 35 for Muse Glimmer against 38 for Qwen3.6 27B and 36 for Kimi K2.5, with Gemma 4 31B at 30 and Llama 4 Maverick at 14. On GDPval-AA v2 Elo it is 953 against Qwen's 1141, which is not a close call. Its Openness Index of 44 ties the best in the field, which is the Apache 2.0 licence showing up as a score.

The number we would actually make a decision on is the hallucination rate: 82%, against 49% for Qwen3.6 27B and 34% for Gemini 3.5 Flash-Lite. Its AA-Omniscience Index is -33. A model that confabulates at that rate is a fine summariser and a bad answer engine, and it is worth knowing before you point it at anything factual. That is a capability judgement rather than a pricing one, but it changes what the $20.00 is buying.

Where this leaves a forecast

Our read. Muse Glimmer is a good model with an uninteresting rental price, and the interesting thing about it is the licence rather than the rate card. At low volume, $0.35 and $1.50 is perfectly fine and not worth optimising; the difference between it and Gemma 4 on a small job is a few dollars a month and you should pick on capability. At high volume the calculation inverts, because there is no second host to negotiate against and the only lever you have is the download.

The specific thing we would not do is write $0.35 into a twelve-month forecast. Not because it is likely to rise, but because a single-host price has no history of being tested and no mechanism forcing it down. Every other open model on the table above got cheaper when a second host turned up. If Fireworks ships the listing it was named for, the number moves. If it does not, the number sits exactly where Together put it, and the only pressure on it comes from people deciding a 24 GB card is cheaper than a subscription.

You can line this up against the rest of the field on the pricing page or run your own token mix through the cost calculator. The 10:1 ratio we used above is a reasonable default for agent work and a bad one for chat, and moving it changes the ordering of that table more than you would expect.

What we checked, and the pages that disagree

  • Meta AI Research: Introducing Muse Glimmer - The August 10, 2026 release date, the Apache 2.0 licence, logit distillation from Muse Spark during pre-training (the blog does not name a Muse Spark version, and an open-weight Muse Spark 1.2 is separately still forthcoming), support for 100+ languages, the twelve-name launch partner list of which nine are local runtimes, and the 3.1x RTX 5090 speedup. Meta's blog gives no absolute tokens-per-second figure: the 74.9 and 233.4 numbers come from the model card below. Also the hero image on this post
  • OpenRouter: meta/muse-glimmer-30b - The endpoints API returns exactly one entry, provider name Together, at $0.35 input and $1.50 output with a 131,072 context and a null max completion tokens. The page states the model is hosted by one provider and that requests are forwarded directly. Its listed creation date of August 9 is one day earlier than Meta's announcement, most likely a staging timestamp
  • Together AI: pricing - The first-party $0.35 and $1.50 card, matching OpenRouter exactly, which is what establishes that OpenRouter is passing the price through rather than setting one
  • Artificial Analysis: Muse Glimmer - Intelligence Index 35, Openness Index 44, GDPval-AA v2 Elo 953, MMMU-Pro 74%, Terminal-Bench v2.1 52%, Tau3-Banking 24%, AA-Omniscience Index -33 and the 82% hallucination rate, plus the peer figures for Qwen3.6 27B (38, 1141 Elo, 49%), Kimi K2.5 (36, 1004), Gemma 4 31B (30) and Llama 4 Maverick (14). It also states Meta is not serving the model on its own API. Note its model page lists a 256K context, which the config and every host contradict
  • Hugging Face: Muse-Glimmer-30B-GGUF - The only directly observable file sizes: 15.9 GB for UD-Q4_K_XL, 19.2 to 21.8 GB at 5-bit, 26.3 GB at 6-bit, 29.6 GB for Q8_0 and 55.7 GB at BF16. These are what the quantisation table above uses, in preference to the 55, 59.55 and roughly 60 GB figures that circulate for full precision
  • Hugging Face: meta-models/Muse-Glimmer-30B - The 131,072 context window, the January 4, 2026 knowledge cutoff, text-plus-image input with text-only output, the ~29.6B total including a ~1.8B vision encoder, the full nine-row benchmark table against Qwen3.6-27B and Gemma4-31B, and the RTX 5090 decode figures of 74.9 tokens per second baseline and 233.4 with DFlash speculative decoding. This is also the card giving TerminalBench 2.1 as 51.7 for Muse Glimmer and 60.7 for Qwen. The 32,768 figure circulating for max output appears in the Hugging Face blog post inside an example agent configuration and is not a model limit
  • NVIDIA NIM: meta-muse-glimmer-30b - Confirms the 131,072 context independently, and that NVIDIA's hosting runs under its API trial terms with no published per-token price, which is why it does not appear as a price in the comparison
  • SiliconANGLE: Meta releases Muse Glimmer - The count of roughly two dozen benchmarks tested and first place in about half of them, which is the check on reading Meta's own nine-row table as a record. Also the framing that this is Meta's first open model in more than a year
  • Fireworks AI: serverless pricing - Checked on August 11, 2026 and Muse Glimmer is absent, despite Fireworks being named as a launch partner. Its published tier for models above 16B is a flat $0.90 per million, which is where the model would land if the listing appears
  • OpenRouter: gpt-oss-120b, Gemma 4 31B, Qwen3.6 27B and Llama 4 Maverick - The multi-host ranges behind the spread table: gpt-oss-120b $0.03 to $0.35 input and $0.17 to $0.95 output, Gemma 4 31B $0.08 to $0.99 and $0.34 to $1.49, Qwen3.6 27B $0.289 to $0.60 and $2.00 to $3.60, Llama 4 Maverick $0.20 to $0.35 and $0.696 to $1.15, Mistral Small 3.2 24B $0.075 to $0.094 and $0.20 to $0.30. Every monthly bill in this post uses the cheapest single host that sells both legs, not the lowest input and lowest output taken from different hosts. For Qwen3.6 27B that is Chutes at $0.30 and $2.00, since the $0.289 minimum input belongs to Morph, whose output is $2.40; for Gemma 4 31B it is OpenInference at $0.08 and $0.35
  • OpenAI: API pricing - GPT-5.6 Luna at $0.20 input and $1.20 output, the standard card following the July 30 cut. OpenRouter currently shows a further 50% discount on Luna and there is a separate long-context tier at $0.40 and $1.80, neither of which is used in the table above