GPT-6 Astra is a OpenAI closed-API frontier model. CompareLLM Coding Reported rating is 1660. Active context window extends to 1.1M tokens. Commercial API token pricing is listed at $10/1M tok input and $50/1M tok output per 1M tokens. OpenAI frontier reasoning model for coding, agents, long-context work, and visual inputs. Numbers below are dated snapshots from empirical benchmark harnesses.
Independent evaluation answering: "Is GPT-6 Astra the right model for your workload & budget?"
Listed output price $50.00/1M tokens. Quality is the category rating above, not this price.
Dated snapshot metrics aggregated from official evaluators and API providers with visual relative score bars.
Our category lists (Coding, Reasoning, Chat, Agents). Named suites such as SWE-bench Pro are recipe inputs, not this rating.
LMArena, SWE-bench Pro, LiveBench and similar suites appear with a source and as-of date. They are not CompareLLM's rating.
How much of the recipe is present. Missing inputs stay missing; they are never filled with 0 or 50.
Compact streaming class. Exact tok/s stays in the fact sheet, not in compact cells.
Compact class for time-to-first-token. Exact milliseconds stay in the fact sheet.
Commercial API price per 1 Million prompt/output tokens (~750k words).
| Benchmark Metric & Meaning | Reported Score & Capability Fill | CompareLLM Review |
|---|---|---|
Input token price Cost to Prompt (Input tokens) | $10/1M tok | Source not recorded · Sep 11, 2026 |
Cached input token price Cost to Reuse Cached Input | $1/1M tok | Source not recorded · Sep 11, 2026 |
Maximum output (tokens) Maximum Response Length | 128k tokens | Source not recorded · Sep 11, 2026 |
Context window (tokens) Context Window Capacity (tokens) | 1.1M tokens | Source not recorded · Sep 11, 2026 |
Output token price Cost to Generate (Output tokens) | $50/1M tok | Source not recorded · Sep 11, 2026 |
Generation speed (tok/s) Writing Speed (tok/s) | <50 tok/s | Source not recorded · Sep 13, 2026 |
Time to first token Waiting Time (before it replies) | 2s+ | Source not recorded · Sep 13, 2026 |
Peers in the same closed API frontier performance tier — not a jump to an unrelated frontier SKU.
Select any rival to launch a side-by-side empirical benchmark comparison with winner deltas.
Compare GPT-6 Astra against
Models may use different benchmarks and test settings. This is an indicative composite, not a controlled head-to-head comparison or community Elo. Admin-approved sentiment estimates fill categories without accepted benchmark results. Estimates are labelled and do not increase benchmark coverage. Coverage refers to the configured recipe, not confidence.
Release recent-models-2026-09-23-r1 · recipe reported-text-2026-09-11-r1 · method reported-with-estimates-v2 · research through 2026-09-23
Overall benchmark coverage includes missing applicable categories. Estimated categories contribute to the rating but add no benchmark coverage.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Repository-level work receives the largest share for coding relevance.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Complement repository work with terminal-based tasks.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. Codex-like developer instructions; see source footnote 8. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. One contribution for related code suites; prefer Extended for breadth, independently of achieved score.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. Dataset revision not specified in launch table. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Broad scientific reasoning is a central signal.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Give equal weight to difficult mathematics rather than letting one discipline dominate.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. Tools enabled; dataset revision not specified. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Tool-assisted academic problems broaden the evidence, with a smaller weight because tools affect outcomes.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. Responses API harness modifications disclosed in source footnote 1; measures the model with its harness. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Interactive abstraction complements academic tasks; small weight limits harness-specific influence.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
Editorial estimate, anchored to the Artificial Analysis Intelligence Index v4.3, on which GPT-6 Astra (max) scores 53, and placed mid-range, because this category has too few measured results to define a band. AA's index is a composite of ten evaluations and is not this category's recipe, so this is an ordering anchor rather than a measurement. Any accepted result that clears the evidence thresholds replaces it.
As of 2026-09-23 · review on 2026-12-23 · Editorial estimate (not a benchmark result)
Sentiment source 1 →Lech Mazur (third-party report) · reported 2026-09-05
Config: High effort; evaluator-v2/v3 comparison evidence with bridge validation; 50-model cohort pinned to source commit. Model judges compare stories written to matching constrained briefs. · Effort: high
Admin-authored relevance policy, not empirical difficulty calibration. Direct creative-writing evaluator evidence; scoped to constrained stories and a frozen comparison cohort.
Published estimated win chance is 92%; separate relative comparison score is 3.5 (interval 3.4–3.6). These units are not interchangeable. Not an absolute writing grade; not a website pairwise prediction.
Missing benchmark families: instruction-compliance.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. End-to-end professional tasks are central to agent usefulness.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. Offline subset and partial credit, not the full online task set. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Desktop interaction is a complementary core agent workload; offline/partial-credit scope is disclosed.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Include multistep workflow automation.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Include browsing and information retrieval without dominating desktop evidence.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
Editorial estimate, anchored to the Artificial Analysis Intelligence Index v4.3, on which GPT-6 Astra (max) scores 53, and placed mid-range, because this category has too few measured results to define a band. AA's index is a composite of ten evaluations and is not this category's recipe, so this is an ordering anchor rather than a measurement. Any accepted result that clears the evidence thresholds replaces it.
As of 2026-09-23 · review on 2026-12-23 · Editorial estimate (not a benchmark result)
Sentiment source 1 →Artificial Analysis (third-party report) · reported 2026-09-03
Config: GPT-6 Astra (max), τ³-Banking panel in the article's Intelligence Evaluations chart. AA describes 97 banking-support tasks, pass@1 averaged over repeats and backend-state grading. Exact historical repeat count/harness revision unspecified. · Effort: max
Authored relevance policy: support outcomes receive 25%. This is the sole present family for Astra, so its Chat rating is provisional and reflects banking-support task completion only, not general conversation quality.
Provisional chat/support proxy, 25% recipe coverage. Visually verified label: 41%. Source chart: https://cdn.sanity.io/images/6vfeftx9/articles/086540355fe9876fc4106c484b6066fd819e8205-2924x4184.png . Banking task success does not measure conversational tone, helpfulness or general satisfaction. Whole-percent source precision retained.
Missing benchmark families: conversation-quality, instruction-compliance.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. No tools; interface localization is a limited part of vision understanding. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Direct no-tool localization gets the largest weight.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. Tool-assisted geometric overlap. Image interpretation plus code generation, not native image generation. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Add spatial interpretation with a smaller share because code/tools contribute.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
OpenAI (vendor-reported) · reported 2026-09-03
Config: Publisher-reported best across effort settings; research/API environment. Exact per-row effort and full run configuration were not disclosed. Optical score recognition, not audio understanding or music generation. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Symbol recognition broadens this recipe; the narrow music-notation domain gets a modest share.
Proposed accepted input in an unpublished draft. Scope/version is limited to the cited launch result; does not assert an exact undisclosed dataset revision.
What this model is available as, what it takes to run, and where its identity comes from.
Available only through the provider's hosted API.
gpt-6-astraopenai/gpt-6-astraConcise empirical overview formatted for citations and prompt context

OpenAI filled out the GPT-6 line on 22 September. Sol keeps GPT-5.6 Sol's $2/$10 price, scores one point higher and streams 37% faster. Luna halves input cost and cuts output cost by 58%, for one point less on the Intelligence Index.

Anthropic's Claude Opus 5.5 scores 58 on the Artificial Analysis Intelligence Index, five points clear of anything else in our catalog, and lists at $4 input and $20 output per million tokens. That is 20% below Opus 5 and 60% below Fable 5.1 and GPT-6 Astra.

Seventeen models from thirteen vendors reached the market in September, and the top of the Intelligence Index moved from 51 to 53. No vendor publishes a roadmap, so here is what the release record actually supports.
Plain-English methodology and leaderboard answers
Elo is a crowd vote on which hidden answer people liked more — not a school test. What is preference Elo? · Methodology
Follow this model in your watchlist, set it as your global comparison baseline, or assign it to your custom production stack.
Track updates & rank changes
Compare all models against this
Assign to custom architecture
Compare side-by-side vs all
Plotted against all active catalog models (50th percentile = catalog median).
Ranks in the top tier (≥75th percentile) for Context, Reasoning.
Pre-computed production rankings across developer workloads based on empirical benchmark capability, throughput, and operational economics.