# GPT-6 Sol

GPT-6 Sol is an AI model released by OpenAI on Sep 22 2026. It has a 1.05M context window. Tracked results include 68.8% on DeepSWE 1.1, 60.5% on OSWorld 2.0 and 56.4% on Agent's Last Exam (pass@1).

## Facts

| Field | Value |
| --- | --- |
| Model | GPT-6 Sol |
| Developer | OpenAI |
| Release date | Tuesday, Sep 22 2026 |
| Licensing | Proprietary |
| Context window | 1.05M |

## API pricing

All rates in USD per 1,000,000 tokens, pay-as-you-go.

| Tier | Input | Cached input | Cache write | Output |
| --- | --- | --- | --- | --- |
| Up to 272K input tokens | $2.00 | $0.20 | $2.50 | $10.00 |
| Over 272K input tokens | $4.00 | $0.40 | $5.00 | $15.00 |

OpenAI also publishes a Fast mode for this model at exactly twice the standard rate in every column — $4.00 per million input tokens and $20.00 per million output at short context.
Verified September 22, 2026 against the first-party source: https://developers.openai.com/api/docs/pricing

## Tracked benchmark scores

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| DeepSWE 1.1 | 68.8% | Lab | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. |
| OSWorld 2.0 | 60.5% | Lab | Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Version 2.0 is a harder, refreshed task set. Higher is better. |
| Agent's Last Exam (pass@1) | 56.4% | Lab | A hard set of desktop and operating-system tasks an AI agent has to finish by looking at the screen and working the machine itself. The score is the share it passes outright — partial credit does not count. Higher is better. |
| AutomationBench | 33.2% | Lab | Tests whether the AI can run real multi-step business workflows — the kind of end-to-end office processes companies want to automate — from start to finish. Higher is better. |

## About GPT-6 Sol

GPT-6 Sol, released September 22, 2026, was the mid-tier model of OpenAI's GPT-6 generation, launched with GPT-6 Luna nineteen days after GPT-6 Astra opened the line and on the same day as Anthropic's Claude Opus 5.5. OpenAI trained both with methods similar to Astra's and pitched them as the cost-efficiency end of the family: Astra stayed "our best model across the board", while Sol and Luna carried its gains in professional work, factuality, coding and computer use to faster, cheaper models. The price was the headline. Sol launched at $2.00 per million input tokens and $10.00 per million output — half the $4.00 and $20.00 OpenAI had charged for GPT-5.6 Sol under promotional pricing, and a fifth of Astra's $10.00 and $50.00 — with cached input at $0.20, cache writes at $2.50, and prompts past 272K input tokens at double on input and 1.5x on output.

The launch post argued the model on cost per task rather than on raw score. At xhigh effort it scored 33.2% on AutomationBench at $0.27 a task, which OpenAI set against 30.3% for GPT-6 Astra at low effort and 26.9% for Claude Opus 5 at max effort, at 9% of Opus 5's cost per task; on Agents' Last Exam it reached 56.4% at max effort, above Opus 5's best at 60% lower cost. In coding it posted 68.8% on DeepSWE v1.1, within 1.1 points of the 69.9% Claude Fable 5 had reached at xhigh effort and at roughly 80% lower cost per task, and OpenAI said it matched Claude Fable 5.1 on FrontierCode 1.1 without publishing the figure. On the offline set of OSWorld 2.0 it scored 60.5% at xhigh effort, level with Claude Opus 5 at medium effort. OpenAI also reported about half as many factual errors as GPT-5.6 Sol on its internal evaluation, approaching Astra's reliability at much lower cost.

The API surface was the 5.6 family's: a 1,050,000-token context window, 128,000 max output tokens, text and image in and text out, and a reasoning-effort control from none through low, medium, high and xhigh to max, with medium as the default, on the Responses, Chat Completions and Batch endpoints. OpenAI shipped improved prompt caching alongside the pair, with a 90% discount on cached reads and cache preserved across changes to reasoning effort and tool availability. Sol's knowledge cutoff was April 20, 2026. It reached ChatGPT Work and Codex on launch day for Plus, Pro, Business, Enterprise and Edu users, and the API as gpt-6-sol.

## Questions and answers

### When was GPT-6 Sol released?

GPT-6 Sol was released by OpenAI on Tuesday, Sep 22 2026.

### Who made GPT-6 Sol?

GPT-6 Sol was built by OpenAI. Creators of ChatGPT and the GPT series of models. Pioneered large-scale language model research.

### How much does GPT-6 Sol cost?

GPT-6 Sol costs $2.00 per million input tokens and $10.00 per million output tokens through the OpenAI API. Cached input is $0.20 per million tokens. Those are the rates for the “Up to 272K input tokens” tier; 1 other pricing tier is published for this model. Rates are pay-as-you-go API prices verified against OpenAI's published pricing on September 22, 2026.

### What benchmark scores did GPT-6 Sol get?

GPT-6 Sol reports 4 tracked benchmark scores — DeepSWE 1.1: 68.8%; OSWorld 2.0: 60.5%; Agent's Last Exam (pass@1): 56.4%; AutomationBench: 33.2%. Tracked scores may come from lab reports or independent benchmarks; source details accompany the benchmark data.

### What is the context window of GPT-6 Sol?

GPT-6 Sol has a context window of 1.05M. That is the maximum amount of input plus output the model can hold in a single request.

### Is GPT-6 Sol open source?

No. GPT-6 Sol is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after GPT-6 Sol?

OpenAI's previous tracked release was GPT-6 Astra on Sep 3 2026, 19 days earlier. It was followed by GPT-6 Luna on Sep 22 2026.


---

Canonical page: https://aireleasetracker.com/model/openai/gpt-6-sol
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
