# Claude Sonnet 5

Claude Sonnet 5 is an AI model released by Anthropic on Jun 30 2026. Tracked results include 80% on BullshitBench v2, 85.2% on SWE-Bench Verified and 57.4% on Humanity's Last Exam (with tools).

## Facts

| Field | Value |
| --- | --- |
| Model | Claude Sonnet 5 |
| Developer | Anthropic |
| Release date | Tuesday, Jun 30 2026 |
| Licensing | Closed |

## API pricing

All rates in USD per 1,000,000 tokens, pay-as-you-go.

| Tier | Input | Cached input | 5 min cache write | 1 hr cache write | Output |
| --- | --- | --- | --- | --- | --- |
| All documented contexts | $2.00 | $0.20 | $2.50 | $4.00 | $10.00 |

Verified August 18, 2026 against the first-party source: https://platform.claude.com/docs/en/about-claude/pricing

## Tracked benchmark scores

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| BullshitBench v2 | 80% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better. |
| Gray Swan IPI (k = 1) | 0.6% | Lab | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. Gray Swan's indirect prompt injection benchmark measures how often such an attack succeeds when the attacker gets a single try. Lower is better. |
| Gray Swan IPI (k = 10) | 4.7% | Lab | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. This variant gives the attacker 10 tries and counts an attack as successful if any of them works. Lower is better. |
| Gray Swan IPI (k = 15) | 5.9% | Lab | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. This variant gives the attacker 15 tries and counts an attack as successful if any of them works. Lower is better. |
| SWE-Bench Pro | 63.2% | Lab | Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better. |
| SWE-Bench Verified | 85.2% | [BenchLM](https://benchlm.ai), retrieved 2026-08-24 | Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better. |
| CursorBench 4.0 | 34.1% | [CursorBench](https://cursor.com/cursorbench), retrieved 2026-09-28 | Cursor's own test of coding agents on ambiguous, multi-file tasks taken from real Cursor sessions — editing, refactoring, investigating a codebase, understanding what the user meant, managing jobs and following a design. Cursor runs each model at several reasoning efforts; each release here carries the score of its best listed effort. Scores aren't comparable with earlier CursorBench versions. Higher is better. |
| FrontierCode v1.1 (Main) (main split) | 42.4% | Lab | A set of very hard, frontier-difficulty coding tasks an AI agent has to complete end to end. The score is the share of tasks in the main split it solves. Higher is better. |
| Next.js Evals | 81% | [Next.js Evals](https://nextjs.org/evals) | Vercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better. |
| Supabase Evals (with skills) | 79.7% | [Supabase Evals](https://supabase.com/evals) | Supabase's own open benchmark: a coding agent is dropped into a real Supabase project and asked to do real work — set up a schema, fix a broken security policy, debug an Edge Function — and every run is checked against a live Supabase stack. This is the headline number, where the agent has Supabase's own skills loaded, as most people building on Supabase would. The score is the share of scenarios it got right. Higher is better. |
| Supabase Evals (no skills) | 75.4% | [Supabase Evals](https://supabase.com/evals) | The same Supabase scenarios, but with none of Supabase's skills loaded — so it measures what the model already knows about building on Supabase, rather than how well it follows Supabase's supplied instructions. Higher is better. |
| Frontier-Bench v0.1 | 14.6% | Lab | A hard, ever-evolving set of real computer tasks — coding, system administration, data work, and more — that an AI agent has to complete on its own. Run by the Harbor / Laude Institute team as the successor to Terminal-Bench (v0.1 is the first release of the task set). The score is the share of tasks solved. Higher is better. |
| Terminal-Bench 4.0 | 12.42% | [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), retrieved 2026-08-30 | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Version 4.0 recalibrated how much time, CPU and memory each task gets, removed eight tasks and fixed nineteen, so fewer runs fail for reasons that have nothing to do with the model. Scores are not comparable with earlier versions. Higher is better. |
| Terminal-Bench 2.1 | 80.4% | Lab | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better. |
| BrowseComp | 84.7% | [BenchLM](https://benchlm.ai), retrieved 2026-07-11 | Can the AI browse the web and track down hard-to-find answers? Higher is better. |
| Humanity's Last Exam (no tools) | 43.2% | Lab | Humanity's Last Exam — extremely hard expert questions across many subjects, written so you can't just look up the answer. “No tools” means the AI answers on its own. Higher is better. |
| Humanity's Last Exam (with tools) | 57.4% | Lab | Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better. |
| OSWorld-Verified | 81.2% | Lab | Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Higher is better. |
| GDPval-AA | 1618 | Lab | Measures how well the AI does economically valuable knowledge work, judged against human experts. Shown as a rating (like a chess Elo) — higher is better. |
| GDPval-AA v2.1 | 1449 | Lab | economically valuable knowledge work (v2.1, Crowd-BT Elo fit) |
| AA-Briefcase v1.1 | 1359 | Lab | Artificial Analysis agentic office-work eval (Elo, v1.1 rating fit) |
| Chartography (no tools) | 15.6% | Lab | The same chart-centred test with no tools: the AI has to read each chart unaided. Scores run far lower than the with-tools version, so read the two as separate tests. Higher is better. |
| threejseval | 1246 | [threejseval](https://threejseval.com/ranking) | Every model gets the same prompt — "the Eiffel Tower", "a glass fishbowl", "a robot arm picking toys into a box" — and builds a 3D scene in Three.js. Real people then see two scenes side by side, names hidden, and vote for the one they prefer. The votes become a chess-style Elo rating on threejseval.com, averaged across all the prompts. It measures whether the scene looks and moves right to a human eye, not whether the code passes a test. Higher is better. |

## Questions and answers

### When was Claude Sonnet 5 released?

Claude Sonnet 5 was released by Anthropic on Tuesday, Jun 30 2026.

### Who made Claude Sonnet 5?

Claude Sonnet 5 was built by Anthropic. AI safety company building the Claude family of models. Founded in 2021 by former OpenAI researchers.

### How much does Claude Sonnet 5 cost?

Claude Sonnet 5 costs $2.00 per million input tokens and $10.00 per million output tokens through the Anthropic API. Cached input is $0.20 per million tokens. Rates are pay-as-you-go API prices verified against Anthropic's published pricing on August 18, 2026.

### What benchmark scores did Claude Sonnet 5 get?

Claude Sonnet 5 reports 23 tracked benchmark scores — BullshitBench v2: 80%; Gray Swan IPI (k = 1): 0.6%; Gray Swan IPI (k = 10): 4.7%; Gray Swan IPI (k = 15): 5.9%; SWE-Bench Pro: 63.2%; SWE-Bench Verified: 85.2%; CursorBench 4.0: 34.1%; FrontierCode v1.1 (Main) (main split): 42.4%; Next.js Evals: 81%; Supabase Evals (with skills): 79.7%; Supabase Evals (no skills): 75.4%; Frontier-Bench v0.1: 14.6%; Terminal-Bench 4.0: 12.42%; Terminal-Bench 2.1: 80.4%; BrowseComp: 84.7%; Humanity's Last Exam (no tools): 43.2%; Humanity's Last Exam (with tools): 57.4%; OSWorld-Verified: 81.2%; GDPval-AA: 1618; GDPval-AA v2.1: 1449; AA-Briefcase v1.1: 1359; Chartography (no tools): 15.6%; threejseval: 1246. Tracked scores may come from lab reports or independent benchmarks; source details accompany the benchmark data.

### Is Claude Sonnet 5 open source?

No. Claude Sonnet 5 is a closed model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after Claude Sonnet 5?

Anthropic's previous tracked release was Claude Mythos 5 on Jun 9 2026, 21 days earlier. It was followed by Claude Opus 5 on Jul 24 2026.


---

Canonical page: https://aireleasetracker.com/model/anthropic/claude-sonnet-5
Site index: https://aireleasetracker.com/llms.txt
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
