Eleven decision models launched in 17 days and none of them bill for output. On a support-ticket classifier output was a quarter of the bill, and Cloudflare's Clef costs more per ticket than GPT-6 Luna.
Eleven of them shipped between September 15 and October 1, and OpenAI announced its own without a price. We put the same support ticket through every rate card we could find.

Photo by Andre Benz on Unsplash
- A classifier writes almost nothing. On GPT-6 Luna output was about 24% of the bill, so free output only takes off that quarter.
- Input rates run from $0.04 to $0.24. At $0.04, Liquid's d1 and Perplexity's Decider come to $0.014 per 1,000 tickets, under a third of Luna's $0.047.
- Cloudflare's Clef lands at $0.085, above Luna with reasoning off and about level with Luna at medium.
What you're buying
A decision model doesn't write. You send it a state (a ticket, a chat turn, an agent's scratchpad) and a list of named questions. Each question is a yes/no, a choice between up to 255 options, or a score on a scale of up to 10 levels. You get back probabilities. TypeSafe started it with Jev on September 15, and most of what followed copies its /v1/systemone request format, so switching between them is mostly a base URL change.
The jobs are the ones teams currently hand to a small chat model with a JSON schema: route this ticket, pick a tool, is this safe to auto-reply, which model should handle this prompt.
| Model | Out | Input / 1M | Context | Weights |
|---|---|---|---|---|
| TypeSafe Jev | Sep 15 | $0.042 | 32K | closed |
| Mapika decider-2b | Sep 16 | self-host | - | Apache 2.0 |
| Together Tev1 4B | Sep 23 | $0.04 | 32K | open, licence unstated |
| Fastino GLiNER2.5-Decide | Sep 24 | not published | - | Apache 2.0 |
| Upstage Solar Decide | Sep 28 | $0.05 | 524K | closed |
| Liquid d1 | Sep 29 | $0.04 | 65K | closed |
| Inception Mercury Decide | Sep 30 | free tier only | 32K | closed |
| Cloudflare Clef-flash | Oct 1 | $0.09 | 65K* | Apache 2.0 |
| Cloudflare Clef | Oct 1 | $0.24 | 65K* | Apache 2.0 |
| Perplexity pplx-decider-v1-27b | Oct 1 | $0.04 | 262K | Apache 2.0 |
| AWS Strands Decider 2B | Oct 1 | self-host | 4K | Apache 2.0 |
Output is $0 on every hosted model here. Prices from vendor docs where published, otherwise the OpenRouter or Vercel AI Gateway listing, read October 5, 2026. Jev accepts 64K per request, 32K for the state plus the longest question. *OpenRouter says Workers AI truncates long text state to roughly the first 2K tokens on both Clef models.
Most of the open ones are Qwen fine-tunes. Clef and Perplexity's Decider both start from Qwen3.8-27B, Clef-flash from Qwen3.5-9B, and Tev1, Strands and Mapika's from Qwen3.5 at 2B to 4B. OpenAI announced a Decisions API at DevDay on September 29. Its developer account says it runs on GPT-6 Luna, but there is still no price, no docs page, and a normal API key gets a 403.
Where the output money actually goes
The best public measurement we found is eesel's ticket test. It couldn't get into the Decisions API, so it ran Luna through the ordinary Responses API with a strict schema: 20 hand-labelled support tickets, each twice, three questions per call (which of 6 queues, which of 3 priorities, safe to auto-reply). Each call used 356 input tokens and, with reasoning off, 23 output tokens.
At Luna's $0.10 and $0.50 that's $0.0000356 of input and $0.0000115 of output per ticket. Output is 24% of the bill. A classifier writes almost nothing, so "we don't charge for output" is a smaller discount than it sounds on a launch slide.
Reasoning is the exception. At medium effort Luna wrote 82 thinking tokens on top of about 25 answer tokens, and output jumped to 60% of an $0.089 bill. If you run your classifier with reasoning on, free output is worth a lot more to you.
The same ticket on every card
We held input at eesel's 356 tokens for everyone, added 23 output tokens to the generative models, and priced 1,000 tickets.
- Claude Haiku 4.5$0.471
- Gemini 3.1 Flash-Lite$0.123
- GPT-6 Luna, medium reasoning$0.0890
- Cloudflare Clef$0.0854
- DeepSeek V4.1 Flash, off-peak$0.0672
- GPT-6 Luna, reasoning off$0.0471
- Gemini 2.5 Flash-Lite$0.0448
- Cloudflare Clef-flash$0.0320
- Upstage Solar Decide$0.0178
- TypeSafe Jev$0.0150
- Liquid d1 / pplx-decider$0.0142
USD per 1,000 tickets. Accent bars are decision models (input only). Grey bars add 23 output tokens at each model's list output price, no caching; the Luna medium row is eesel's measured figure. Our arithmetic from list prices read October 5, 2026.
The $0.04 group lands at about a third of Luna. Clef-flash is a bit under 70% of it. Clef is the odd one: at $0.24 an input token, its input alone costs more than Luna's whole ticket. Luna only becomes the dearer of the two once its output passes 28% of the input token count, which on this ticket means about 100 output tokens. That's reasoning territory.
Two generative rows are worth a second look. Gemini 2.5 Flash-Lite at $0.10 and $0.40 is a hair under Luna and has been on sale for over a year. Claude Haiku 4.5 is 10x Luna and 33x d1 on this job, which is a lot to pay for a three-way label.
One thing we can't control for: a decision request isn't a chat prompt, so the token count won't be exactly 356. The question and option text count as input on Perplexity, Liquid and TypeSafe. TypeSafe's own docs example, a one-line state with a single yes/no question, bills 296 input tokens, so there is overhead per request even when your text is short. Check usage.input_tokens on your own tickets before you trust any row above to the cent.
At these prices, the bill is rarely the reason
Ten million tickets a month is $142 on d1, $150 on Jev, $471 on Luna and $854 on Clef. Most teams don't have ten million tickets. So for most of you the choice comes down to speed and whether the answers are right, and that's where these models differ by more than 3x.
eesel's Luna calls took a median 1.48 seconds with reasoning off and 2.34 at medium. Cloudflare measured Clef-flash at 38.8 milliseconds median across 43 benchmarks, Clef at 209 and Jev at 524, over the network in Jev's case. Strands Decider 2B runs at a 115 ms median on a single RTX 3090. If the decision sits in front of a user, that gap matters more than any column in the price table.
On accuracy, the only independent board is Hugging Face's Jev Decision Index: Jev 57.91, pplx-decider-v1-27b 56.4, Tev1 4B 29.24, Mapika's decider-2b 28.97 and GLiNER2.5-Decide 11.21. Clef, d1 and Strands aren't on it yet. Cloudflare's own numbers have Clef ahead of Jev on BFCL (98.47 against 95.75), API-Bank and BANKING77, and behind it on When2Call (72.37 against 80.97) and BRIGHT. Those are vendor numbers, and the 2K truncation means none of them say much about long states.
Luna got the queue right on 40 of 40 eesel calls and all three answers right on 29. Its weak spot was the auto-reply question, with 7 wrong calls out of 40. If a decision model can't beat that on your tickets, a third of $0.047 isn't a saving.
What we'd try first
For a router or ticket triage with short inputs, Perplexity's Decider at $0.04. It scores about a point and a half behind Jev on the independent index, costs a little less, has a 262K window and the weights are Apache 2.0, so you aren't stuck with one host. Jev is the reference everyone benchmarks against and is worth keeping as the control.
Clef-flash if latency is the whole point and your states are short. Full Clef we'd skip at $0.24 unless its benchmark lead holds up on your data, because the money it costs buys Luna with reasoning on.
And keep an eye on OpenAI. If the Decisions API ships at Luna's $0.10 input with output free, it lands at about $0.036 per 1,000 of these tickets, between Clef-flash and Luna. That is our guess, not OpenAI's price. We covered Luna's card in detail in our GPT-6 Luna post, and DeepSeek V4.1 Flash's peak hours double its row in the chart.
What we checked against
- TypeSafe: Jev models and pricing and launch post - $0.042 input, free output, context, request format
- Cloudflare: Clef decision models and Workers AI model page - prices, bases, latency and benchmark tables
- Perplexity: Decisions quickstart and model card - $0.04 input, limits, token counting
- Liquid AI: decision models - d1 API and what input tokens include
- Together AI: pricing - Tev1 4B at $0.04 (OpenRouter lists $0.042)
- OpenRouter: model listings - d1, Solar Decide and Mercury Decide prices and dates; Clef truncation note
- AWS Strands Labs: Strands Decider - latency on RTX 3090, licence
- Hugging Face: Jev Decision Index - independent scores
- eesel: OpenAI Decisions API pricing - ticket test method, token counts, Luna accuracy and latency
- OpenAI: GPT-6 Luna, Gemini API pricing, Claude pricing and DeepSeek pricing - generative model rates