# GPT-5.3-Codex

GPT-5.3-Codex is an AI model released by OpenAI on Feb 5 2026. At release it scored 24% on BullshitBench v2, 85% on SWE-Bench Verified and 64.7% on OSWorld-Verified.

## Facts

| Field | Value |
| --- | --- |
| Model | GPT-5.3-Codex |
| Developer | OpenAI |
| Release date | Thursday, Feb 5 2026 |
| Licensing | Proprietary |

## API pricing

All rates in USD per 1,000,000 tokens, pay-as-you-go.

| Tier | Input | Cached input | Output |
| --- | --- | --- | --- |
| All documented contexts | $1.75 | $0.175 | $14.00 |

Verified August 18, 2026 against the first-party source: https://developers.openai.com/api/docs/models/gpt-5.3-codex

## Benchmark scores published at release

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| BullshitBench v2 | 24% | Lab | Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better. |
| SWE-Bench Verified | 85% | [BenchLM](https://benchlm.ai), retrieved 2026-08-24 | Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better. |
| Next.js Evals | 81% | Lab | Vercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better. |
| OSWorld-Verified | 64.7% | [BenchLM](https://benchlm.ai), retrieved 2026-08-31 | Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Higher is better. |

## Questions and answers

### When was GPT-5.3-Codex released?

GPT-5.3-Codex was released by OpenAI on Thursday, Feb 5 2026.

### Who made GPT-5.3-Codex?

GPT-5.3-Codex was built by OpenAI. Creators of ChatGPT and the GPT series of models. Pioneered large-scale language model research.

### How much does GPT-5.3-Codex cost?

GPT-5.3-Codex costs $1.75 per million input tokens and $14.00 per million output tokens through the OpenAI API. Cached input is $0.175 per million tokens. Rates are pay-as-you-go API prices verified against OpenAI's published pricing on August 18, 2026.

### What benchmark scores did GPT-5.3-Codex get?

GPT-5.3-Codex reports 4 tracked benchmark scores — BullshitBench v2: 24%; SWE-Bench Verified: 85%; Next.js Evals: 81%; OSWorld-Verified: 64.7%. Scores are the figures published at release by OpenAI.

### Is GPT-5.3-Codex open source?

No. GPT-5.3-Codex is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after GPT-5.3-Codex?

OpenAI's previous tracked release was GPT-5.2 on Dec 11 2025, 56 days earlier. It was followed by GPT-5.3-Codex-Spark on Feb 12 2026.


---

Canonical page: https://aireleasetracker.com/model/openai/gpt-5.3-codex
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
