# Claude Opus 5

Claude Opus 5 is an AI model released by Anthropic on Jul 24 2026. It has a 1M context window. At release it scored 96% on SWE-Bench Verified, 53.4% on FrontierCode v1.1 (Main) (main split) and 43.3% on Frontier-Bench v0.1.

## Facts

| Field | Value |
| --- | --- |
| Model | Claude Opus 5 |
| Developer | Anthropic |
| Release date | Friday, Jul 24 2026 |
| Licensing | Proprietary |
| Context window | 1M |

## API pricing

All rates in USD per 1,000,000 tokens, pay-as-you-go.

| Tier | Input | Cached input | 5 min cache write | 1 hr cache write | Output |
| --- | --- | --- | --- | --- | --- |
| All documented contexts | $5.00 | $0.50 | $6.25 | $10.00 | $25.00 |

Verified August 18, 2026 against the first-party source: https://platform.claude.com/docs/en/about-claude/pricing

## Benchmark scores published at release

| Benchmark | Score | What it measures |
| --- | --- | --- |
| BullshitBench v2 | 73% | Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better. |
| Gray Swan IPI (k = 1) | 0.2% | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. Gray Swan's indirect prompt injection benchmark measures how often such an attack succeeds when the attacker gets a single try. Lower is better. |
| Gray Swan IPI (k = 10) | 1.6% | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. This variant gives the attacker 10 tries and counts an attack as successful if any of them works. Lower is better. |
| Gray Swan IPI (k = 15) | 2% | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. This variant gives the attacker 15 tries and counts an attack as successful if any of them works. Lower is better. |
| SWE-Bench Verified | 96% | Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better. |
| CursorBench v3.2 | 70% | Cursor's own test of harder, real-world coding tasks inside a code editor, on the refreshed v3.2 task set. Scores aren't comparable with v3.1. Higher is better. |
| DeepSWE 1.1 | 68.8% | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. |
| FrontierCode v1.1 (Main) (main split) | 53.4% | A set of very hard, frontier-difficulty coding tasks an AI agent has to complete end to end. The score is the share of tasks in the main split it solves. Higher is better. |
| Next.js Evals | 88% | Vercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better. |
| Supabase Evals (with skills) | 95.5% | Supabase's own open benchmark: a coding agent is dropped into a real Supabase project and asked to do real work — set up a schema, fix a broken security policy, debug an Edge Function — and every run is checked against a live Supabase stack. This is the headline number, where the agent has Supabase's own skills loaded, as most people building on Supabase would. The score is the share of scenarios it got right. Higher is better. |
| Supabase Evals (no skills) | 90.9% | The same Supabase scenarios, but with none of Supabase's skills loaded — so it measures what the model already knows about building on Supabase, rather than how well it follows Supabase's supplied instructions. Higher is better. |
| Frontier-Bench v0.1 | 43.3% | A hard, ever-evolving set of real computer tasks — coding, system administration, data work, and more — that an AI agent has to complete on its own. Run by the Harbor / Laude Institute team as the successor to Terminal-Bench (v0.1 is the first release of the task set). The score is the share of tasks solved. Higher is better. |
| BrowseComp | 90.8% | Can the AI browse the web and track down hard-to-find answers? Higher is better. |
| Humanity's Last Exam (no tools) | 56.3% | Humanity's Last Exam — extremely hard expert questions across many subjects, written so you can't just look up the answer. “No tools” means the AI answers on its own. Higher is better. |
| Humanity's Last Exam (with tools) | 64.7% | Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better. |
| ARC-AGI-3 | 30.2% | The third generation of the ARC-AGI series: instead of static puzzles, the AI is dropped into small interactive game-like environments it has never seen and has to figure out the rules and solve them on its own. Higher is better. |
| ARC-AGI-2 | 90.4% | Puzzle-style tests of abstract reasoning and pattern-finding — the kind of thing people find easy but AIs often struggle with. Higher is better. |
| BioMysteryBench (hard) | 49.4% | Real unsolved-style biology puzzles — the AI has to reason its way to an answer the way a research biologist would. The “hard” split contains the toughest cases. Higher is better. |
| BioMysteryBench (human solved) | 90.1% | Real biology puzzles that human experts have managed to crack — can the AI reach the same answers? Higher is better. |
| OSWorld 2.0 | 70.6% | Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Version 2.0 is a harder, refreshed task set. Higher is better. |
| AutomationBench | 26% | Tests whether the AI can run real multi-step business workflows — the kind of end-to-end office processes companies want to automate — from start to finish. Higher is better. |
| Harvey's Legal Agent Benchmark (Held-out) | 11.7% | Harvey's test of whether an AI agent can complete real legal work, scored on a held-out set of tasks the model makers never see — making the numbers harder to game. Higher is better. |
| HealthBench Professional | 59.8% | Realistic health conversations graded against detailed rubrics written by physicians — can the AI respond the way a careful medical professional would? Higher is better. |
| GDPval-AA v2 | 1861 | economically valuable knowledge work (v2, re-based Elo) |
| Arena Elo (Text) | 1495 | Real people chat with two anonymous AIs side by side and vote for the answer they prefer. Votes become a chess-style Elo rating on arena.ai — it measures which AI people actually like, not test scores. Higher is better. |
| Arena Elo (Code) | 1663 | Like the text arena, but people vote on which AI writes better code. The votes become a chess-style Elo rating on arena.ai. Higher is better. |

## About Claude Opus 5

Claude Opus 5, released July 24, 2026, brought Anthropic's Claude 5 generation to the Opus tier six weeks after Claude Fable 5 opened it — at half Fable's price, keeping Opus 4.8's $5 per million input tokens and $25 per million output. The pitch was flagship-class agentic capability at workhorse pricing: at launch it scored 43.3% on Frontier-Bench v0.1, more than double Opus 4.8's 21.1% and nearly ten points clear of Fable 5, and posted a GDPval-AA v2 Elo of 1861 for knowledge work, the best published score at the time.

The launch card leaned on breadth: 90.8% on BrowseComp for agentic search, 70.6% on OSWorld 2.0 computer use, 64.7% on Humanity's Last Exam with tools, and 30.2% on ARC-AGI-3 — which Anthropic reported as roughly three times the next best published result on the novel problem-solving benchmark. Anthropic made it the default model on Claude Max and the strongest model available on Claude Pro, positioning Opus 5 as the everyday frontier model while Fable 5 kept the edge on a handful of evaluations, including DeepSWE agentic coding and Harvey's held-out legal benchmark.

## Questions and answers

### When was Claude Opus 5 released?

Claude Opus 5 was released by Anthropic on Friday, Jul 24 2026.

### Who made Claude Opus 5?

Claude Opus 5 was built by Anthropic. AI safety company building the Claude family of models. Founded in 2021 by former OpenAI researchers.

### How much does Claude Opus 5 cost?

Claude Opus 5 costs $5.00 per million input tokens and $25.00 per million output tokens through the Anthropic API. Cached input is $0.50 per million tokens. Rates are pay-as-you-go API prices verified against Anthropic's published pricing on August 18, 2026.

### What benchmark scores did Claude Opus 5 get?

Claude Opus 5 reports 26 tracked benchmark scores — BullshitBench v2: 73%; Gray Swan IPI (k = 1): 0.2%; Gray Swan IPI (k = 10): 1.6%; Gray Swan IPI (k = 15): 2%; SWE-Bench Verified: 96%; CursorBench v3.2: 70%; DeepSWE 1.1: 68.8%; FrontierCode v1.1 (Main) (main split): 53.4%; Next.js Evals: 88%; Supabase Evals (with skills): 95.5%; Supabase Evals (no skills): 90.9%; Frontier-Bench v0.1: 43.3%; BrowseComp: 90.8%; Humanity's Last Exam (no tools): 56.3%; Humanity's Last Exam (with tools): 64.7%; ARC-AGI-3: 30.2%; ARC-AGI-2: 90.4%; BioMysteryBench (hard): 49.4%; BioMysteryBench (human solved): 90.1%; OSWorld 2.0: 70.6%; AutomationBench: 26%; Harvey's Legal Agent Benchmark (Held-out): 11.7%; HealthBench Professional: 59.8%; GDPval-AA v2: 1861; Arena Elo (Text): 1495; Arena Elo (Code): 1663. Scores are the figures published at release by Anthropic. It holds the best score among all models tracked here on Gray Swan IPI (k = 1), Gray Swan IPI (k = 10), Gray Swan IPI (k = 15), SWE-Bench Verified, FrontierCode v1.1 (Main) (main split), Frontier-Bench v0.1, Humanity's Last Exam (no tools), Humanity's Last Exam (with tools), ARC-AGI-3, BioMysteryBench (hard), BioMysteryBench (human solved), OSWorld 2.0, Harvey's Legal Agent Benchmark (Held-out), HealthBench Professional and GDPval-AA v2.

### What is the context window of Claude Opus 5?

Claude Opus 5 has a context window of 1M. That is the maximum amount of input plus output the model can hold in a single request.

### Is Claude Opus 5 open source?

No. Claude Opus 5 is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after Claude Opus 5?

Anthropic's previous tracked release was Claude Sonnet 5 on Jun 30 2026, 24 days earlier. It is the most recent Anthropic model tracked on AI Release Tracker.


---

Canonical page: https://aireleasetracker.com/model/anthropic/claude-opus-5
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
