23:34 09 October 2026
GPT-6 Astra is OpenAI's flagship model in this generation: shipped on 2026-09-03, priced at $10 per million input tokens and $50 per million output, advertised with a 1.05M-token context window and a 128K output ceiling, and described by its own vendor as the model for "complex reasoning and coding". It is the earliest of the four models in this comparison to ship, the most expensive per token, and the best-measured of them — 45 of 68 benchmark rows across 7 of 8 categories on the independent coverage board, read 2026-10-07. The GPT-6 Astra API is carried here at list rate with no markup, so everything below can be checked against an invoice rather than a slide.
Astra is the most expensive model here per token and the *second cheapest* per finished task. That inversion is the first thing worth understanding about it.
Astra is one rung of a three-rung line, not a standalone product. The vendor's own routing advice sends complex reasoning and coding to it, cost-sensitive high-volume work to GPT-6 Luna below it, and everything in between to GPT-6.1 Sol in the middle. Read as a family, the line is a price ladder with a stated default, and Astra is the rung the vendor points hard problems at.
It is also a new name rather than a version bump. The independent record's deprecation chains are explicit about replacements — `gpt-6-sol` resolves to `gpt-6.1-sol`, `claude-opus-5` to `claude-opus-5-5`, `grok-4-6` to `grok-4-7` — and Astra carries no such entry, which is what a fresh name in a family looks like rather than a successor to one specific model.
One caveat on the date, and it is ours rather than OpenAI's: our own catalog lists Astra's release as 2026-09-04, one day after the vendor's changelog. When two records disagree by a day, take the vendor's.
Most of Astra's specification sheet reads as a superset of the model below it, which is the point of a flagship.
Three lines in that table are worth a second look.
The context row is the first. The vendor advertises 1.05M and the independent board records a round 1,000,000 — the same pair of numbers Sol has, and the same pattern the board shows for several models, where it reports a vendor's advertised figure as-is in some cases and a round bound in others. Two sources, one of them slightly conservative, and no practical difference: both are large enough that the limit is not what will stop most workloads.
The second is the pricing tier at the bottom, and it is ours rather than the vendor's. On our own rate card, Astra requests above 272,000 input tokens are billed at $20 in and $75 out rather than $10 and $50. That step is OrcaRouter's, not a line from OpenAI's published card, and it only matters to workloads that habitually send very long inputs.
The third is the ladder, because it comes with a restriction. Astra's reasoning settings run from low to max with no `none` rung, and the model is served with fixed sampling — temperature, top_p and logprobs are not exposed as adjustable parameters. If your pipeline depends on sampling a prompt several times at different temperatures and voting, that dependency has to move to a different model or a different design.
Astra is the best-evidenced of the four models here, which is a claim about the size of the record rather than about the scores in it.
On the independent coverage board read 2026-10-07, it publishes 45 of 68 benchmark rows across 7 of 8 categories, with a composite score of 84.9 and a board rank of 2. The one model above it holds 53 of 73 rows and 8 of 8 categories — enough headroom that the ordering between them is not close.
On the independent evaluation harness, also read 2026-10-07 and measured with the model running at its Max reasoning setting:
Two of those rows deserve calling out. The GPQA figure of 96.06% is the only one on the board for any of the four subjects — the other three carry no row there at all. And the Intelligence Index of 52.67 sits one point above the mid-tier model in the same family and nearly six below the highest score of the ten models on that harness, so "flagship" here means the top of one vendor's line rather than the top of the board.
The vendor's positioning is unusually explicit about the division of labour, and it is worth quoting because it is the cheapest routing advice available:
"If you're not sure where to start, use GPT-6 Astra, our flagship model for complex reasoning and coding. Choose GPT-6.1 Sol to balance intelligence and cost, or GPT-6 Luna for cost-sensitive, high-volume workloads."
Three things follow from that sentence. First, Astra is defined by workload shape — reasoning and coding — rather than by being newest. Second, the vendor's own pitch for the model below it is *"near-Astra performance"*, which means Astra's results are the reference the rest of the family is measured against rather than an optional upgrade. Third, a lab that expected everyone to buy the flagship would not write a sentence whose middle clause redirects most readers one rung down.
Here is where the per-token price stops telling the story.
Astra charges five times Sol's rate on both directions and finishes a task in roughly three-quarters of the time, but it generates considerably fewer output tokens to get there, and the two effects very nearly cancel: $3.26 a task against Sol's $0.72 is a 4.5× bill, not the 5× the rate card implies. That is still a large multiple. It is a smaller one than the card advertises, and it is the number to budget from, because it already includes how much the model chooses to talk.
Note the effort settings in that table before comparing anything across rows. Astra and Sol are measured at Max, the top rung; Grok 4.7 is measured at Xhigh, one rung lower. A speed or cost comparison that stacks them without that note is comparing three different configurations.
GPT-6 Astra is the top rung of a three-model line: shipped 2026-09-03, $10 in and $50 out, a 1.05M advertised context against a 1,000,000 recorded one, a 128K output ceiling, a reasoning ladder that starts at low rather than `none`, and fixed sampling with no temperature or top_p to set. It carries the thickest independent record of the four models here and the second-highest score on the harness.
The plain-terms answer to whether that makes it the right model is: only if your work is the kind the vendor built the rung for. Astra's case rests on hard reasoning and coding, and on a per-task price that comes in below what its rate card implies. For ordinary high-volume traffic it is four and a half times the bill for a sub-point difference on the Intelligence Index — and the vendor says so itself in the routing sentence.
OrcaRouter carries GPT-6 Astra at list rate on the same key as Sol and Luna, so moving a route up or down the ladder is a model-string change rather than a second procurement.
Sourcing note: the $10 / $50 rates, the cache read and write prices, the 128K output ceiling, the 1.05M advertised context, the low-to-max reasoning ladder and the fixed-sampling restriction are OpenAI's own published claims and have not been independently reproduced. The 272K-token tier ($20 / $75) is OrcaRouter's own rate card, not OpenAI's. All Intelligence Index, GPQA, HLE, SciCode, MMMU-Pro, Terminal-Bench 4.0, LCR, omniscience and GDPval figures, the per-task costs, stream rates and per-task times are from Artificial Analysis' live model pages, read 2026-10-07; GPT-6 Astra, GPT-6.1 Sol and Claude Opus 5.5 are measured at the Max reasoning setting and Grok 4.7 at Xhigh, so no cross-model speed or cost claim should be read from those rows without that caveat. Coverage counts (45 of 68, 7 of 8, 53 of 73) and the composite score 84.9 are from benchlm.ai, read 2026-10-07. The release date 2026-09-03 is the vendor's; our catalog's 2026-09-04 is noted where it appears. All checked 2026-10-07; re-check before 2026-11-01.