# GPT-5.2

GPT-5.2 is an AI model released by OpenAI on Dec 11 2025. At release it scored 38% on BullshitBench v2, 92.4% on GPQA Diamond and 80% on SWE-Bench Verified.

## Facts

| Field | Value |
| --- | --- |
| Model | GPT-5.2 |
| Developer | OpenAI |
| Release date | Thursday, Dec 11 2025 |
| Licensing | Proprietary |

## API pricing

All rates in USD per 1,000,000 tokens, pay-as-you-go.

| Tier | Input | Cached input | Output |
| --- | --- | --- | --- |
| All documented contexts | $1.75 | $0.175 | $14.00 |

Verified August 18, 2026 against the first-party source: https://developers.openai.com/api/docs/models/gpt-5.2

## Benchmark scores published at release

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| BullshitBench v2 | 38% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better. |
| SWE-Bench Verified | 80% | Lab | Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better. |
| BrowseComp | 65.8% | [BenchLM](https://benchlm.ai), retrieved 2026-08-31 | Can the AI browse the web and track down hard-to-find answers? Higher is better. |
| ARC-AGI-2 | 52.9% | [BenchLM](https://benchlm.ai), retrieved 2026-07-11 | Puzzle-style tests of abstract reasoning and pattern-finding — the kind of thing people find easy but AIs often struggle with. Higher is better. |
| GPQA Diamond | 92.4% | Lab | Graduate-level science questions in biology, physics, and chemistry — hard enough that subject-matter PhDs score around 65%. Higher is better. |
| OSWorld-Verified | 47.3% | [BenchLM](https://benchlm.ai), retrieved 2026-08-31 | Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Higher is better. |

## About GPT-5.2

GPT-5.2, released December 11, 2025, closed out OpenAI's 2025 with its strongest general model of the year: 92.4% on GPQA Diamond and 80.0% on SWE-Bench Verified, arriving less than a month after GPT-5.1 and directly answering Google's Gemini 3.0 Pro and Anthropic's Claude Opus 4.5 from the preceding weeks.

It also scored 52.9% on ARC-AGI-2, at the time among the best published results on the abstract-reasoning benchmark. GPT-5.2 anchored OpenAI's lineup through the winter while the coding-focused GPT-5.3-Codex branch shipped in February 2026, and was ultimately succeeded by GPT-5.4 in March 2026.

## Questions and answers

### When was GPT-5.2 released?

GPT-5.2 was released by OpenAI on Thursday, Dec 11 2025.

### Who made GPT-5.2?

GPT-5.2 was built by OpenAI. Creators of ChatGPT and the GPT series of models. Pioneered large-scale language model research.

### How much does GPT-5.2 cost?

GPT-5.2 costs $1.75 per million input tokens and $14.00 per million output tokens through the OpenAI API. Cached input is $0.175 per million tokens. Rates are pay-as-you-go API prices verified against OpenAI's published pricing on August 18, 2026.

### What benchmark scores did GPT-5.2 get?

GPT-5.2 reports 6 tracked benchmark scores — BullshitBench v2: 38%; SWE-Bench Verified: 80%; BrowseComp: 65.8%; ARC-AGI-2: 52.9%; GPQA Diamond: 92.4%; OSWorld-Verified: 47.3%. Scores are the figures published at release by OpenAI.

### Is GPT-5.2 open source?

No. GPT-5.2 is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after GPT-5.2?

OpenAI's previous tracked release was GPT-5.1-Codex-Max on Nov 19 2025, 22 days earlier. It was followed by GPT-5.3-Codex on Feb 5 2026.


---

Canonical page: https://aireleasetracker.com/model/openai/gpt-5.2
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
