# GPT-6 Astra

GPT-6 Astra is an AI model released by OpenAI on Sep 3 2026. It has a 1.05M context window. At release it scored 96% on GPQA Diamond, 64.5% on FrontierCode v1.1 (Extended) (extended split) and 67 on AA Coding Agent Index.

## Facts

| Field | Value |
| --- | --- |
| Model | GPT-6 Astra |
| Developer | OpenAI |
| Release date | Thursday, Sep 3 2026 |
| Licensing | Proprietary |
| Context window | 1.05M |

## API pricing

All rates in USD per 1,000,000 tokens, pay-as-you-go.

| Tier | Input | Cached input | Cache write | Output |
| --- | --- | --- | --- | --- |
| Up to 272K input tokens | $10.00 | $1.00 | $12.50 | $50.00 |
| Over 272K input tokens | $20.00 | $2.00 | $25.00 | $75.00 |

OpenAI also publishes a Fast mode for this model at exactly twice the standard rate in every column — $20.00 per million input tokens and $100.00 per million output at short context.
Verified September 3, 2026 against the first-party source: https://developers.openai.com/api/docs/pricing

## Benchmark scores published at release

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| Auto-review circumvention (Internal) | 0% | Lab | OpenAI's internal safety check on how often a model finds ways around its own automated review — the guardrail that inspects what it is about to do. This one counts failures, so lower is better and zero is the goal. |
| DeepSWE 1.1 | 74.1% | Lab | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. |
| FrontierCode v1.1 (Main) (main split) | 53.3% | Lab | A set of very hard, frontier-difficulty coding tasks an AI agent has to complete end to end. The score is the share of tasks in the main split it solves. Higher is better. |
| FrontierCode v1.1 (Extended) (extended split) | 64.5% | Lab | frontier-difficulty agentic coding tasks (v1.1, extended split) |
| AA Coding Agent Index | 67 | Lab | Artificial Analysis' overall score for coding agents, combining three coding benchmarks with what each run costs and how many tokens it burns. It rates a model paired with a particular agent harness rather than the model alone, so the same model scores differently in different tools. Higher is better. |
| Database Migration Tasks (OpenAI Internal) | 63.9% | Lab | OpenAI's own test of moving a database from one schema or system to another without breaking what depends on it — the migration work that has to be right the first time. Higher is better. |
| BenchCAD | 95.9% | Lab | Can the AI do mechanical design as code? Given a drawing or a description of an industrial part — a gear, a spring, a drill bit — it has to write or edit the parametric CAD program that builds it, and the program is run to check the shape really comes out right. Higher is better. |
| Terminal-Bench 4.0 | 57.9% | Lab | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Version 4.0 recalibrated how much time, CPU and memory each task gets, removed eight tasks and fixed nineteen, so fewer runs fail for reasons that have nothing to do with the model. Scores are not comparable with earlier versions. Higher is better. |
| Terminal-Bench-Science 0.1 | 64.6% | Lab | The same command-line setup as Terminal-Bench, pointed at scientific work: the AI has to drive research tooling and computational workflows through to a result, rather than administer a machine. Version 0.1 is the first release of the task set, and scores run lower than on the general board. Higher is better. |
| BrowseComp | 91.5% | Lab | Can the AI browse the web and track down hard-to-find answers? Higher is better. |
| ExploitBench | 100% | Lab | A 'capability ladder' for security research, built by CMU researchers: the AI is given known bugs in Chrome's V8 engine and scored on how far it gets toward a working exploit inside a research sandbox — from understanding the patch to triggering a crash. Higher is better. |
| ARC-AGI-3 | 99.9% | Lab | The third generation of the ARC-AGI series: instead of static puzzles, the AI is dropped into small interactive game-like environments it has never seen and has to figure out the rules and solve them on its own. Higher is better. |
| FrontierMath (Tier 4 (v2)) | 97.6% | Lab | The rebuilt edition of FrontierMath's hardest tier — research-level maths of the kind professional mathematicians work on. It is a different question set from the first Tier 4, and scores on it run far higher, so read the two as separate tests rather than progress. Higher is better. |
| GeneBench-Pro | 39% | Lab | Real genomics and biomedical analyses done end to end: the AI gets a messy dataset and a question, and has to work through the chain of statistical decisions to a verifiable answer — the job a computational biologist does before a research or clinical decision gets made. Higher is better. |
| MedChemBench (Internal) | 49.7% | Lab | OpenAI's own drug-discovery test: reading chemical structures, predicting how potent or toxic a compound will be, choosing between candidate molecules, and planning a synthesis route — the everyday judgement calls of a medicinal chemist. Higher is better. |
| GPQA Diamond | 96% | Lab | Graduate-level science questions in biology, physics, and chemistry — hard enough that subject-matter PhDs score around 65%. Higher is better. |
| Agent's Last Exam (pass@1) | 59.3% | Lab | A hard set of desktop and operating-system tasks an AI agent has to finish by looking at the screen and working the machine itself. The score is the share it passes outright — partial credit does not count. Higher is better. |
| SRE-Bench | 99.2% | Lab | Can the AI keep production running? It is dropped into a broken Kubernetes system and has to diagnose the incident and fix it safely, the way an on-call site-reliability engineer would. Scored here on the best of four attempts. Higher is better. |
| AutomationBench | 41.4% | Lab | Tests whether the AI can run real multi-step business workflows — the kind of end-to-end office processes companies want to automate — from start to finish. Higher is better. |
| HealthBench Professional (length-adjusted) | 63.4% | Lab | The same physician-graded health conversations as HealthBench Professional, scored with an adjustment for how long the answer is — so a model cannot gain by padding its reply. The adjustment moves scores by several points, so these numbers are not interchangeable with the unadjusted ones. Higher is better. |
| AA Intelligence Index | 61.2 | Lab | Artificial Analysis composite intelligence index across evals |
| Design Tasks (OpenAI Internal) | 50% | Lab | OpenAI's own set of professional design briefs, scored on whether the finished work is what a designer would have handed over. Higher is better. |
| Data Science Tasks (OpenAI Internal) | 40.9% | Lab | OpenAI's own set of data-science jobs — taking a dataset and a question through cleaning, analysis and a defensible answer. Higher is better. |
| OpenScore String Quartets | 0.84 | Lab | Can the AI read sheet music? It is shown scanned pages of string quartets and has to transcribe the notation — pitches, beams, accidentals and all — with the score measuring how close the transcription lands to the real thing. Runs 0 to 1, and higher is better. |

## About GPT-6 Astra

GPT-6 Astra, released September 3, 2026, opened OpenAI's GPT-6 generation and was the company's first new general-purpose flagship since the three-model GPT-5.6 family — Sol, Terra and Luna — arrived on June 26, 2026. OpenAI had trailed the model for a month before launch without ever announcing it as a product: on August 1, 2026 it named Astra "our next major model" in a research post, reporting that an internal version had produced new results on ten long-standing open mathematical problems, each with a machine-checkable proof in the Lean theorem prover. Capability demonstrated in a paper before a model existed to sell was an unusual way to introduce a flagship, and it set the terms for the four weeks that followed.

The launch table led on reasoning: 99.9% on ARC-AGI-3, against the 30.2% Claude Opus 5 had posted on the same interactive-environment board in July, 97.6% on the rebuilt second edition of FrontierMath Tier 4, and 96.0% on GPQA Diamond. Agentic work moved by similar margins — 64.6% on Terminal-Bench Science 0.1 and 41.4% on AutomationBench, both more than double GPT-5.6 Sol's figures in the same run, and 74.1% on DeepSWE v1.1. The security rows are the ones that explain the wait: 100.0% on ExploitBench, the Carnegie Mellon capability ladder that scores how far a model gets turning a known V8 bug into a working exploit, and 99.2% on SRE-Bench over four attempts.

The professional and coding tables ran closer. Astra took Terminal-Bench 4.0 at 57.9% against Claude Fable 5.1's 55.8%, BrowseComp at 91.5%, and 0.84 on the OpenScore String Quartets sheet-music transcription set where GPT-5.6 Sol had managed 0.19 — but on FrontierCode 1.1 its 53.3% on the main split and 64.5% on the extended set both sat fractionally behind the Claude models OpenAI ran beside it, and it trailed on Artificial Analysis' two composite indices: 61.2 on the Intelligence Index against Fable 5.1's 65.7, and 67.0 on the Coding Agent Index. For a generational release the shape of the table was lopsided — enormous margins on reasoning, security and the internal professional sets, and a coding field the previous generation had already crowded.

That capability is what OpenAI spent August evaluating. On August 7 the company said it had paused parts of the work on Astra because preliminary results could not rule out that the model reached the Critical cybersecurity level in its Preparedness Framework — the threshold describing a model able to find and carry out attacks against well-defended systems without human help, and the first time OpenAI had put one of its own models in that bracket. It returned to the question on September 1 in "Path to Astra: critical capabilities and frontier safeguards", saying the further evaluations were complete and the safeguards it had built around the model were sufficient to ship. Astra launched two days later.

## Questions and answers

### When was GPT-6 Astra released?

GPT-6 Astra was released by OpenAI on Thursday, Sep 3 2026.

### Who made GPT-6 Astra?

GPT-6 Astra was built by OpenAI. Creators of ChatGPT and the GPT series of models. Pioneered large-scale language model research.

### How much does GPT-6 Astra cost?

GPT-6 Astra costs $10.00 per million input tokens and $50.00 per million output tokens through the OpenAI API. Cached input is $1.00 per million tokens. Those are the rates for the “Up to 272K input tokens” tier; 1 other pricing tier is published for this model. Rates are pay-as-you-go API prices verified against OpenAI's published pricing on September 3, 2026.

### What benchmark scores did GPT-6 Astra get?

GPT-6 Astra reports 24 tracked benchmark scores — Auto-review circumvention (Internal): 0%; DeepSWE 1.1: 74.1%; FrontierCode v1.1 (Main) (main split): 53.3%; FrontierCode v1.1 (Extended) (extended split): 64.5%; AA Coding Agent Index: 67; Database Migration Tasks (OpenAI Internal): 63.9%; BenchCAD: 95.9%; Terminal-Bench 4.0: 57.9%; Terminal-Bench-Science 0.1: 64.6%; BrowseComp: 91.5%; ExploitBench: 100%; ARC-AGI-3: 99.9%; FrontierMath (Tier 4 (v2)): 97.6%; GeneBench-Pro: 39%; MedChemBench (Internal): 49.7%; GPQA Diamond: 96%; Agent's Last Exam (pass@1): 59.3%; SRE-Bench: 99.2%; AutomationBench: 41.4%; HealthBench Professional (length-adjusted): 63.4%; AA Intelligence Index: 61.2; Design Tasks (OpenAI Internal): 50%; Data Science Tasks (OpenAI Internal): 40.9%; OpenScore String Quartets: 0.84. Scores are the figures published at release by OpenAI. It holds the best score among all models tracked here on Auto-review circumvention (Internal), FrontierCode v1.1 (Extended) (extended split), AA Coding Agent Index, Database Migration Tasks (OpenAI Internal), BenchCAD, Terminal-Bench-Science 0.1, ExploitBench, ARC-AGI-3, FrontierMath (Tier 4 (v2)), GeneBench-Pro, MedChemBench (Internal), GPQA Diamond, Agent's Last Exam (pass@1), SRE-Bench, HealthBench Professional (length-adjusted), AA Intelligence Index, Design Tasks (OpenAI Internal), Data Science Tasks (OpenAI Internal) and OpenScore String Quartets.

### What is the context window of GPT-6 Astra?

GPT-6 Astra has a context window of 1.05M. That is the maximum amount of input plus output the model can hold in a single request.

### Is GPT-6 Astra open source?

No. GPT-6 Astra is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after GPT-6 Astra?

OpenAI's previous tracked release was GPT-5.6-Cyber on Aug 10 2026, 24 days earlier. It is the most recent OpenAI model tracked on AI Release Tracker.


---

Canonical page: https://aireleasetracker.com/model/openai/gpt-6-astra
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
