Skip to main content
TokenCost logoTokenCost
Model ReleaseSeptember 4, 2026ยท13 min read

GPT-6 Astra's rate card is Claude Fable 5's rate card, to the cent, on every line both companies publish. Anthropic moved one of those lines two days before Astra shipped, and it is now the only line where the two flagships disagree.

OpenAI released GPT-6 Astra on September 3 at $10.00 per million input tokens, $1.00 cached, $12.50 for a cache write and $50.00 output. Those are the four numbers Claude Fable 5 has carried since launch. On September 1, Anthropic shipped Fable 5.1 and cut exactly one of them: the cache read, from $1.00 to $0.25. So the newest flagship in the market lines up with the one Anthropic just retired, and the gap between the two live flagships is one cell wide.

Two identical concrete doorways in a weathered wall, one open to darkness, one sealed with a steel door

Photo by Kiyota Sage on Unsplash

Five things to know before the tables

  • Seven things tie between Astra and Fable 5.1: input, output, the five-minute cache write, both batch legs, the 128,000-token output ceiling, and the 1.25x multiplier each house uses to price a cache write. The cache read is the whole rate-card difference, $1.00 against $0.25.
  • That gap exists because Anthropic broke a formula it had never broken before. Every Claude model ever sold priced a cache hit at one tenth of base input. Fable 5.1 and Mythos 5.1 are the first two at 0.025x, and the exception lives in a footnote rather than the table.
  • Astra sits on the old formula: ten dollars input, one dollar cached, 0.100x on the nose, alongside GPT-5.6 Sol, GPT-5.6 Cyber, GPT-5.3 Codex, Opus 5, Sonnet 5 and Haiku 4.5. OpenAI did not answer Anthropic's cut. It priced its flagship on the convention both companies shared until Tuesday.
  • The second difference is a threshold rather than a rate. Any Astra request over 272,000 input tokens reprices the entire request at 2x input and 1.5x output, and Fable 5.1 bills its full million-token window flat. Below that line Astra costs about 13% more on a cache-heavy agent loop. Above it, the same two cards produce a 2.46x gap.
  • Which leaves the interesting question. OpenAI's launch post claims lower cost per task than Fable 5.1 on three benchmarks, by 31%, 63% and 86%. On a card that ties on input and output, a cost claim cannot be a price claim. It is arithmetic about token counts wearing a dollar sign, and it back-solves cleanly, which we do below.

The two cards, and the one Anthropic left behind

Fable 5 is still on Anthropic's price list, sitting in the main seventeen-row table with no retirement note against it, unlike Opus 4.1, Opus 4, Sonnet 4 and Haiku 3.5, which are all marked retired. It is the models overview page, not the price list, that files it under legacy models still available. Read the three columns across: Astra matches the superseded card, not the current one, everywhere OpenAI publishes a comparable figure.

Line, per million tokensGPT-6 AstraFable 5.1Fable 5, legacy
Input$10.00$10.00$10.00
Output$50.00$50.00$50.00
Cache write, 5-minute$12.50$12.50$12.50
Cache write, 1-hournot published$20.00$20.00
Cache read$1.00$0.25$1.00
Batch input$5.00$5.00$5.00
Batch output$25.00$25.00$25.00
Batch cache read$0.50$0.125$0.50
Above 272,000 input tokens$20.00 / $2.00 / $75.00no surchargeno surcharge
Context window1,050,0001,000,0001,000,000
Max output128,000128,000128,000

Two of those rows deserve a caveat rather than a comparison. OpenAI publishes no one-hour cache tier for Astra, so that row is empty rather than equal, and Anthropic's $20.00 has nothing to sit opposite. And the 272,000 row is the one place where the shape of the two products differs rather than the price: Astra's surcharge applies to the whole request, not to the tokens past the line, so it is a step function rather than a tier.

Both houses charged a tenth for a cache hit. One of them stopped.

The 0.1x cache-read multiplier was not a coincidence between two companies, it was a convention. Here is every current flagship and near-flagship on both price lists, and the multiplier each one implies.

ModelVendorInputCache readMultiple
GPT-6 AstraOpenAI$10.00$1.000.100x
GPT-5.6 CyberOpenAI$12.50$1.250.100x
GPT-5.6 SolOpenAI$4.00$0.400.100x
GPT-5.3 CodexOpenAI$1.75$0.1750.100x
Claude Fable 5Anthropic$10.00$1.000.100x
Claude Opus 5Anthropic$5.00$0.500.100x
Claude Sonnet 5Anthropic$2.00$0.200.100x
Claude Haiku 4.5Anthropic$1.00$0.100.100x
Claude Fable 5.1Anthropic$10.00$0.250.025x
Claude Mythos 5.1Anthropic$10.00$0.250.025x

Eight rows at 0.100x, two at 0.025x, and the two exceptions are three days old. Cache writes tell the same story from the other side: both companies charge 1.25x base input for a five-minute write, which is why Astra and Fable 5.1 land on the identical $12.50 without either of them copying the other. Two firms arrived at the same three multipliers from the same base price. Then one of them moved, and the model that shipped 48 hours later did not follow.

We wrote up the Fable 5.1 cut on the day it landed, when the comparison was Fable 5.1 against its own predecessor. Astra turns it into a cross-vendor comparison without changing a single number.

What 4x on one line is actually worth

A 4x gap sounds decisive until you notice what it is 4x of. Take a thousand agent turns, each sending a 60,000-token prompt and getting 4,000 tokens back, and sweep the cache hit rate. Output is $50.00 on both models and never moves, so it anchors the whole comparison.

Cache hit rateGPT-6 AstraFable 5.1Astra premium
0%$800.00$800.000.0%
50%$530.00$507.504.4%
70%$422.00$390.508.1%
85%$341.00$302.7512.6%
90%$314.00$273.5014.8%
95%$287.00$244.2517.5%
99%$265.40$220.8520.2%

At zero caching the two models are the same price, because at zero caching the only line that differs is not used. At a 99% hit rate, about as good as a well-built agent loop gets, Astra costs 20.2% more. That is the practical ceiling on this workload shape, and it is a long way below 4x. The reason is arithmetic: your saving equals 75% of whatever share of the bill was cache reads, and at an 85% hit rate with 4,000 output tokens per turn, cache reads are only $51.00 of Astra's $341.00, which is 15.0%. Three quarters of 15.0% is 11.2%, and $341.00 minus 11.2% is $302.75.

Flip that around and the gap only approaches 4x as output approaches zero and the hit rate approaches one, which describes a classification job, not an agent. If your product is a long-context retrieval endpoint that returns fifty tokens, the cache-read line is most of your bill and Anthropic's $0.25 is close to a 4x saving. If your product writes code, output dominates and the two models are within about 13% of each other. Neither of those is the headline number, and neither vendor published either of them.

One asymmetry survives the batch discount intact. Both companies halve everything for batch, which halves the cache read too, so Astra is $0.50 against Fable 5.1's $0.125 and the ratio stays exactly 4.00x. There is no tier on either side that closes it.

The second difference is a number, not a price: 272,000

OpenAI's model page states it in one sentence: prompts with more than 272,000 input tokens are priced at 2x input and cache rates and 1.5x output for the full request. Not for the excess. For the request. So the cost of a job does not rise smoothly as prompts grow, it doubles at a point. Here is the same thousand-turn loop at 85% cache hits, sized to sit exactly on the line and then one token past it, with Fable 5.1 alongside for scale.

272,000-token promptsAstra, at the lineAstra, one token overFable 5.1, either
Cached input, 231.2M tokens$231.20$462.40$57.80
Uncached input, 40.8M tokens$408.00$816.00$408.00
Output, 4.0M tokens$200.00$300.00$200.00
Total per 1,000 turns$839.20$1,578.40$665.80

One token adds $739.20 across a thousand turns, an 88.1% increase on a prompt that grew by 0.0004%. Against Fable 5.1 the same job goes from 1.26x to 2.37x on the strength of that single token, and nothing in Anthropic's pricing responds, because Anthropic has no threshold to respond with. Every one of Fable 5.1's million tokens bills at the same rate as its first.

The other thing worth noticing is where 272,000 sits inside the window OpenAI advertises. Astra's context is 1,050,000 tokens, of which 922,000 is the input ceiling and 128,000 the output ceiling, and those two add to 1,050,000 exactly, so the headline figure is a sum of two limits rather than a single addressable space. Measured against the 922,000 you can actually fill with a prompt, the threshold sits at 29.5%. Roughly 70% of Astra's usable input range bills at double. Fable 5.1 advertises 1,000,000, all of it at one price, and the same 128,000 output ceiling.

Push the prompts to 400,000 tokens and the two cards stop looking alike at all.

1,000 turns, 400,000-token prompts, 85% cachedGPT-6 AstraClaude Fable 5.1
Cached input, 340.0M tokens$680.00$85.00
Uncached input, 60.0M tokens$1,200.00$600.00
Output, 4.0M tokens$300.00$200.00
Total per 1,000 turns$2,180.00$885.00

Nearly $1,300 apart on a job that costs $885.00 to run on the other card, from two rate cards that agree on input and output. The long-context rule does all of it. If you are building anything that routinely carries a large document set or a long agent history in the prompt, this threshold matters more than any per-token figure on either page, and it is the reason we would not describe these two models as similarly priced without saying how long your prompts are.

Inside OpenAI's own lineup, Astra is Sol times two and a half

Ten lines, one multiplier, no exceptions. GPT-5.6 Sol has been at $4.00 and $20.00 since its August price cut, and Astra is that card scaled by exactly 2.5 everywhere, including the derived tiers.

Line, per million tokensGPT-5.6 SolGPT-6 AstraMultiple
Input$4.00$10.002.50x
Cached input$0.40$1.002.50x
Cache write$5.00$12.502.50x
Output$20.00$50.002.50x
Batch and Flex input$2.00$5.002.50x
Batch and Flex output$10.00$25.002.50x
Fast input$8.00$20.002.50x
Fast output$40.00$100.002.50x
Long-context input$8.00$20.002.50x
Long-context output$30.00$75.002.50x

Both models put the long-context threshold at 272,000 and both apply the same 2x input, 1.5x output rule, so the 2.50x holds at any prompt length, any cache hit rate and any service tier. That is unusually convenient for anyone modelling a migration: there is no crossover point to find, no mix where Astra wins on rate. If a task takes fewer than 40% as many tokens on Astra as on Sol, it is cheaper. Otherwise it is not. One number decides the whole thing.

With one caveat about the denominator. Sol's $4.00 and $20.00 are promotional: OpenAI's model page says the rate is available at least through November 21, 2026, which is a floor on when the price could move rather than a date on which it will. If Sol ever returns to the $5.00 and $30.00 it carried before August, the ratio falls to 2.00x on input and 1.67x on output and stops being a single number. Astra carries no promotional label and no expiry date at all.

One tier does not survive the scaling. Fast mode is 2x on both models, but OpenAI's pricing page states that Fast is unavailable for GPT-6 Astra under EU data residency, and directs those requests to Standard. Regional processing endpoints also carry a 10% uplift for models released recently enough to qualify, which is the same 1.1x Anthropic charges for US-only inference and the same 10% Google charges on non-global Vertex endpoints. Three vendors, one number, and on Astra it stacks on top of a card that may already have doubled.

Every cost claim in the launch post is a token claim

OpenAI attaches a cost figure to five benchmark results, phrased as estimated API cost per task. Because the rate cards involved are either a fixed multiple of Astra's (Sol, at 0.4x on every line) or identical to it (Fable 5.1, on input and output), those cost figures invert exactly. Cost equals rate times tokens; if the rate ratio is known and constant, the token ratio falls straight out.

BenchmarkCompared withOpenAI's cost claimRate ratioImplied tokens, rival vs Astra
Terminal-Bench 4.0GPT-5.6 Sol9% lower2.50x2.75x
Terminal-Bench 4.0Claude Fable 5.163% lower1.00x2.70x or more
Terminal-Bench Science 0.1Claude Fable 5.131% lower1.00x1.45x or more
Terminal-Bench Science 0.1, low-cost settingGPT-5.6 Sol27% lower2.50x3.42x
BenchCAD, with toolsGPT-5.6 Sol43% lower2.50x4.39x
DeepSWE v1.1GPT-5.6 Sol57% lower2.50x5.81x
BenchCAD, with toolsClaude Fable 5.186% lower1.00x7.14x or more
GPQA Diamond, low-cost settingGPT-5.6 Sol37% lower2.50x3.97x

The Sol rows are exact. Sol's card is 0.4x Astra's on input, cached input and output alike, so the ratio survives any token mix and any cache hit rate. Terminal-Bench 4.0 at 9% lower cost means Astra spends 36.4% of Sol's tokens on the same task. The Fable 5.1 rows carry a floor rather than an equals sign, because Astra's cache read is 4x Fable 5.1's: if OpenAI's estimate includes cached tokens at all, Astra was paying more for them, so its token count has to be even lower than the naive inversion suggests. Hence 63% lower cost on a tied card implies Fable 5.1 spending at least 2.70x Astra's tokens.

Read down the last column and the spread is the finding. Against Fable 5.1 the implied token ratio runs from 1.45x on Terminal-Bench Science to 7.14x on BenchCAD, a factor of five between two benchmarks in the same blog post. Against Sol it runs 2.75x to 5.81x. That is not necessarily anyone misleading anyone: token efficiency genuinely is task-specific, and OpenAI says these figures come from the configurations shown. But it does mean there is no such thing as "Astra is 63% cheaper". There is only Astra being 63% cheaper on Terminal-Bench 4.0, in a setup OpenAI chose, using OpenAI's estimate of a competitor's bill.

The sixth claim is the honest version of the same thing and skips the dollar sign entirely: on Agents' Last Exam, OpenAI says Astra uses approximately 65% fewer output tokens than Claude Opus 5. Opus 5 bills $25.00 output against Astra's $50.00, so 0.35 times the tokens at twice the rate is 0.70 times the output cost. A 30% saving, on a model that is twice the price per token. That is the shape of every one of these comparisons, stated plainly.

The scores themselves are not in dispute

Worth separating from the cost arithmetic: OpenAI puts a Claude Fable 5.1 column on roughly a dozen rows, and the two we can check against Anthropic's own launch material, 55.8% on Terminal-Bench 4.0 and 52.6% on Terminal-Bench Science 0.1, are the numbers Anthropic published for its own model three days ago. Nobody is arguing about the scores. The unauditable part is the cost estimate sitting next to them.

BenchmarkGPT-6 AstraGPT-5.6 SolClaude Fable 5.1
Terminal-Bench 4.057.9%37.3%55.8%
Terminal-Bench Science 0.164.6%22.4%52.6%
BenchCAD, with tools95.9%83.3%84.3%
DeepSWE v1.174.1%72.7%not quoted
GPQA Diamond96.0%94.6%93.7%
FrontierMath Tier 4 (v2)97.6%83.0%87.8%
ARC-AGI-399.9%7.8%not quoted
OSWorld 2.0, latency simulation72.6%65.7%not quoted
ExploitBench100%78.5%70%
Exploit Gym42.4%30.3%30.4%
SRE-Bench, single attempt88.0%55.9%12.5%

Two rows need their footnotes carried over. The Terminal-Bench Science figure for Sol, 22.4%, is described by OpenAI as Sol's best result and is compared against Astra at a deliberately lower-cost setting scoring 61.1%, not against the 64.6% in the table. GPQA works the same way: Astra's headline is 96.0%, and the 37% cost saving is quoted against a different, cheaper Astra configuration scoring 94.9%. Cost claims and top-line scores come from different runs throughout the post, which is disclosed, and easy to lose.

One score in that table comes with an outside voice attached. Greg Kamradt of the ARC Prize Foundation is quoted in the launch post saying Astra surpassed their human action-efficiency baseline on 96% of ARC-AGI-3 levels, which is a claim about how efficiently the model learns rather than how much it costs. It is the one figure in the post attributed to someone who does not sell the model, and it still is not a price.

The cyber results carry a caveat of a different kind. ExploitBench at 100% and ExploitGym at 42.4% were measured without production safeguards, on a model OpenAI says meets the Critical cybersecurity threshold under its Preparedness Framework. The Astra you can buy refuses the tasks those benchmarks measure. It is a published score for a configuration that is not for sale, which is the same problem we hit with Gemini 3.8 Flash Cyber two days ago, from the opposite direction: there, a claim about cost with no score attached.

Ten dollars buys a million tokens. It does not buy the same text.

There is a hole under every cross-vendor price comparison, this one included, and Anthropic's own documentation is where it is easiest to see. Claude 4.7 and later models use a newer tokenizer that Anthropic says produces approximately 30% more tokens for the same text than the one Sonnet 4.6 and earlier used, with the exact increase depending on content. Fable 5.1 is on the newer tokenizer. So $10.00 per million tokens on Fable 5.1 and $10.00 per million tokens on Fable 5 were never the same price per page, and Anthropic says so in a note under its own table.

That 30% is Anthropic measuring itself against itself, not against OpenAI, and neither company publishes a conversion between their tokenizers. We are not going to invent one. What the disclosure does establish is the size of the effect: a tokenizer change inside a single vendor moved the real price of a document by roughly a third while every number on the rate card stayed still. A tokenizer difference across two vendors can be at least that large in either direction, and no rate card anywhere captures it. Which is worth holding in mind before treating $10.00 and $10.00 as a tie, and worth holding in mind again when reading OpenAI's token-count claims about a competitor whose tokens are not its own.

You cannot buy this model, which is why nobody downstream has priced it

OpenAI's post says Astra reaches developers through the OpenAI API, Microsoft Azure and AWS Bedrock. As of today only the first of those has a number. Azure's OpenAI pricing page carries an undated banner saying GPT-6 prices are in processing for publication and directing readers to the blog for them, which is a pointer rather than a rate. Bedrock has published nothing. OpenRouter, which usually lists a new frontier model within hours, had no Astra row at all when we pulled its full catalogue on September 4: 427 models, zero matches on either "astra" or "gpt-6". It does carry Fable 5.1, at $10.00, $50.00, $0.25 cache read, $12.50 and $20.00 cache writes, and a batch variant with every line halved including the cache read at $0.125. Anthropic's card is fully mirrored downstream; OpenAI's does not exist downstream yet.

That is because you cannot buy Astra. OpenAI says it is rolling out to enterprises in the Trusted Access Program, with API access and the consumer and business plans following in the coming days. A published rate card for a gated model is not a price you can pay; it is a price you can plan against. Fable 5.1 went generally available on the Claude API, Bedrock, Google Cloud and Microsoft Foundry on day one, so for the moment the comparison in this post is between a model you can call this afternoon and one you can put in a spreadsheet.

When the API does open, the rate limits are the next thing to read. Astra's input ceiling is 922,000 tokens, and tokens-per-minute budgets are small enough that the top of the context window is out of reach for most accounts.

Usage tierRequests per minuteTokens per minuteMax-length prompts per minute
Tier 1500500,000no
Tier 25,0001,000,000yes, one
Tier 35,0002,000,000yes, two
Tier 410,0004,000,000yes, four
Tier 515,00040,000,000yes, forty-three

OpenAI publishes Tier 1 through Tier 5 for this model and no free tier. A Tier 1 account cannot send a single maximum-length request at all: 922,000 tokens is 84% more than the entire 500,000-token minute it is allowed. Tier 2 fits exactly one, with 78,000 tokens to spare. The long-context surcharge is priced across a range that most accounts cannot reach in the first place, which is a strange thing to say about a headline feature, and it means the 272,000 threshold is the operative number for far more people than the 1,050,000 one.

Three numbers of yours settle it, in order of how much they move

Since the two cards tie on input and output, the rate card cannot decide anything for you and neither can OpenAI's benchmark table. Three of your own numbers can.

The first is your median prompt length, and specifically whether it crosses 272,000 tokens. That single number is worth more than everything else on this page: below it, Astra runs about 13% above Fable 5.1 on a cache-heavy loop, and above it, 2.46x. If your prompts live at 250,000 tokens and grow, you have a cliff coming that no amount of caching will soften, and the mitigation is prompt design rather than tier selection.

Next, work out the share of your bill that is cache reads. That is the only line where the two flagships differ, and your saving from Anthropic's $0.25 is three quarters of that share and nothing more. Pull it from an invoice, not from a blog post. Most agent workloads we have priced land between 10% and 25%, which puts the Fable 5.1 advantage between 8% and 19% on the current cards.

Last, and hardest, count tokens per finished task on your own work. Every cost argument OpenAI made is a claim about this number, and the implied values in its own post disagree with each other by a factor of five depending on the benchmark. That spread is the reason to measure rather than extrapolate: on your tasks the answer could plausibly land anywhere in it, and a model at 2.5x the rate of Sol needs to be under 40% of Sol's token spend just to break even. Run the two models on fifty real tasks, count the tokens, and price it yourself. Our calculator takes token counts and a cache hit rate directly, and the model catalogue carries every line quoted above with its source.

One prediction we will not make: whether Astra's $1.00 cache read stays there. OpenAI has cut a flagship price mid-generation once already this year, taking Sol down 20% on input and 33% on output in August. Anthropic has now shown that the cache read is a line you can move on its own without touching anything else. Those two facts sit close together, and the model that would answer the question is not for sale yet.

Sources, and the gaps we left open

  • OpenAI: GPT-6 Astra model page - Read September 4, 2026. The $10.00, $1.00, $12.50 and $50.00 card; the sentence putting the long-context threshold at more than 272,000 input tokens with 2x input and cache rates and 1.5x output across the full request; cache writes at 1.25x uncached input; Batch and Flex at 50% and Fast at 2x; 1,050,000 context, 128,000 max output, April 30, 2026 knowledge cutoff; the reasoning.effort ladder of low, medium, high, xhigh and max; the Trusted Access Program rollout note; and the per-tier rate limit table
  • OpenAI: API pricing - The full Astra grid across Standard, Batch, Flex and Fast in both context bands, including $20.00 / $2.00 / $75.00 long-context Standard and $20.00 / $100.00 short-context Fast; the GPT-5.6 Sol and Cyber rows used for the 2.50x table; the note that Fast mode is unavailable for GPT-6 Astra under EU data residency; and the 10% uplift on regional processing endpoints
  • OpenAI: GPT-6 Astra, a new generation of intelligence - September 3, 2026. Every benchmark score and every cost claim back-solved above, in OpenAI's own wording: 9% and 63% lower estimated API cost per task on Terminal-Bench 4.0 against Sol and Fable 5.1, 31% and 27% on Terminal-Bench Science, 43% and 86% on BenchCAD, 57% on DeepSWE v1.1 against Sol, 37% on GPQA Diamond at a lower-cost setting, and approximately 65% fewer output tokens than Claude Opus 5 on Agents' Last Exam. Also the rollout sentence naming the OpenAI API, Microsoft Azure and AWS Bedrock, and the statement that Astra meets the Critical cybersecurity threshold under the Preparedness Framework while refusing the tasks the cyber benchmarks measure
  • Anthropic: Claude pricing - The seventeen-row model table giving Fable 5.1 and Fable 5 at $10.00 base, $12.50 and $20.00 cache writes, $50.00 output, and $0.25 against $1.00 on cache hits; footnote 1 stating that Fable 5.1 and Mythos 5.1 price cache hits at 0.025x base input where all other models use the standard 0.1x; the note that Claude 4.7 and later use a tokenizer producing approximately 30% more tokens for the same text; the 1.1x US inference geography multiplier; and the 10% premium on Bedrock and Google Cloud regional endpoints
  • Anthropic: models overview - Fable 5.1 at a 1,000,000-token context window and 128,000 max output, adaptive thinking always on at high effort by default, a June 2026 reliable knowledge cutoff, and the API ID claude-fable-5-1. No long-context tier appears on this page or the pricing page, which is the basis for describing the window as flat
  • OpenRouter: full model catalogue - Pulled September 4, 2026. 427 models, no row matching "astra" or "gpt-6". Fable 5.1 present at $10.00 / $50.00 with input_cache_read 0.00000025, input_cache_write 0.0000125 and input_cache_write_1h 0.00002 per token, plus a batch variant with every field halved
  • TokenCost: the Fable 5.1 cache read cut and the GPT-5.6 Sol price cut - Our September 2 and August posts establishing the two prior cards this one is measured against
  • What we could not establish. Neither AWS nor Microsoft has published an Astra rate; Azure's OpenAI pricing page carries an undated banner saying GPT-6 prices are pending, and Bedrock has nothing, so no cloud comparison appears here. OpenAI publishes no one-hour cache tier for Astra, so that row is blank rather than zero. No cross-vendor tokenizer conversion exists between o200k_base and Anthropic's current tokenizer, so the tokenizer section states the size of Anthropic's self-reported effect and stops there. The implied token ratios are arithmetic on OpenAI's own published percentages and inherit whatever assumptions sit behind the phrase "estimated API cost"; OpenAI does not publish the token counts themselves, the harnesses, or the reasoning effort used for the competitor runs. Astra's launch date is given as September 3, 2026 from the announcement; the model page says rolling out today without dating itself. Two FrontierMath figures circulate: the launch post's prose says Astra saturates FrontierMath Tier 4 with a 98% score, while its comparison table gives 97.6% for the v2 set, and the table row is the one quoted above because it is the one with a version attached. OSWorld 2.0 carries no Fable 5.1 column, and the 70.2% on that row belongs to Claude Opus 5, which several write-ups have misattributed. And the three workload tables are our arithmetic on published rates, not measurements: the token counts and hit rates in them are assumptions, chosen to be legible, and your own numbers are the ones that matter