Claude Sonnet 5
Released
Claude Sonnet 5 is an AI model released by Anthropic on Tuesday, Jun 30 2026, 21 days after Claude Mythos 5. Benchmark results (shown below) cover BullshitBench v2, Gray Swan IPI, SWE-Bench Pro, SWE-Bench Verified, CursorBench 4.0, FrontierCode v1.1 (Main), and 13 more.
Get Anthropic releases and AI news.
One email, every other Monday.
API pricing
- InputWhat you pay for everything you send the model — your question, plus any documents or earlier conversation you include with it.
- $2.00
- Cached inputA reduced rate for text you send over and over. If every request starts with the same instructions or the same document, the provider keeps a copy ready and charges less to read it again.
- $0.20
- 5 min cache writeA one-off charge for storing text so later requests can reuse it at the cheaper cached rate. This option keeps it for five minutes.
- $2.50
- 1 hr cache writeA one-off charge for storing text so later requests can reuse it at the cheaper cached rate. This option keeps it for an hour, so it costs more than the five-minute one.
- $4.00
- OutputWhat you pay for the text the model writes back. It is normally the dearer half: producing an answer costs more than reading one.
- $10.00
Available from
| Amazon Bedrock | global | $2.00 | $10.00 | 1M |
|---|---|---|---|---|
| Anthropic | — | $2.00 | $10.00 | 1M |
| Azure | global | $2.00 | $10.00 | 1M |
| Claude Platform on AWS | — | $2.00 | $10.00 | 1M |
| global | $2.00 | $10.00 | 1M | |
| Amazon Bedrock | eu-west-1 | $2.20 | $11.00 | 1M |
| Amazon Bedrock | us-east-1 | $2.20 | $11.00 | 1M |
| Azure | us | $2.20 | $11.00 | 1M |
| eu | $2.20 | $11.00 | 1M | |
| us | $2.20 | $11.00 | 1M |
Benchmarks
Coding
SWE-Bench ProAgentic coding — Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better.
63.2%
#8 of 23Best: Claude Fable 5.1 · 81.2%
SWE-Bench VerifiedCoding — Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better.via BenchLM
85.2%
#6 of 57Best: Claude Opus 5 · 96%
CursorBench 4.0Agentic coding — Cursor's own test of coding agents on ambiguous, multi-file tasks taken from real Cursor sessions — editing, refactoring, investigating a codebase, understanding what the user meant, managing jobs and following a design. Cursor runs each model at several reasoning efforts; each release here carries the score of its best listed effort. Scores aren't comparable with earlier CursorBench versions. Higher is better.via CursorBench
34.1%
#12 of 13Best: Claude Opus 5.5 · 57.8%
FrontierCode v1.1 (Main)Agentic coding — A set of very hard, frontier-difficulty coding tasks an AI agent has to complete end to end. The score is the share of tasks in the main split it solves. Higher is better.
42.4%main split
#7 of 7Best: Claude Opus 5.5 · 54.4%
Next.js EvalsNext.js coding — Vercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better.
81%
#9 of 32Best: Claude Fable 5.1 · 97%
Supabase EvalsSupabase coding — Supabase's own open benchmark: a coding agent is dropped into a real Supabase project and asked to do real work — set up a schema, fix a broken security policy, debug an Edge Function — and every run is checked against a live Supabase stack. This is the headline number, where the agent has Supabase's own skills loaded, as most people building on Supabase would. The score is the share of scenarios it got right. Higher is better.
79.7%with skills
75.4%no skills
with skills
#7 of 11Best: Claude Opus 5 · 92.8%
no skills
#7 of 11Best: Claude Opus 5.5 · 91.3%
Terminal & CLI
Frontier-Bench v0.1Agentic computer work — A hard, ever-evolving set of real computer tasks — coding, system administration, data work, and more — that an AI agent has to complete on its own. Run by the Harbor / Laude Institute team as the successor to Terminal-Bench (v0.1 is the first release of the task set). The score is the share of tasks solved. Higher is better.
14.6%
#7 of 9Best: Claude Opus 5 · 43.3%
Terminal-Bench 4.0Agentic terminal coding — Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Version 4.0 recalibrated how much time, CPU and memory each task gets, removed eight tasks and fixed nineteen, so fewer runs fail for reasons that have nothing to do with the model. Scores are not comparable with earlier versions. Higher is better.via Terminal-Bench
12.42%
#17 of 19Best: Claude Sonnet 5.5 · 70.6%
Terminal-Bench 2.1Agentic terminal coding — Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better.
80.4%
#19 of 32Best: DeepSeek-V4.1-Flash · 90.6%
Agentic & tool use
BrowseCompWeb browsing — Can the AI browse the web and track down hard-to-find answers? Higher is better.via BenchLM
84.7%
#11 of 28Best: GPT-5.6 Sol · 92.2%
OSWorld-VerifiedAgentic computer use — Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Higher is better.
81.2%
#7 of 28Best: Qwen3.8-Max · 86.1%
Reasoning & science
Humanity's Last ExamMultidisciplinary reasoning — Humanity's Last Exam — extremely hard expert questions across many subjects, written so you can't just look up the answer. “No tools” means the AI answers on its own. Higher is better.
43.2%no tools
57.4%with tools
no tools
#7 of 22Best: Claude Fable 5.1 · 60.9%
with tools
#13 of 45Best: Claude Opus 5.5 · 67.7%
Knowledge work
GDPval-AAKnowledge work — Measures how well the AI does economically valuable knowledge work, judged against human experts. Shown as a rating (like a chess Elo) — higher is better.
1618
#7 of 9Best: Claude Fable 5 · 1932
GDPval-AA v2.1Knowledge work — economically valuable knowledge work (v2.1, Crowd-BT Elo fit)
1449
#6 of 6Best: Claude Opus 5.5 · 1846
AA-Briefcase v1.1Knowledge work — Artificial Analysis agentic office-work eval (Elo, v1.1 rating fit)
1359
#4 of 4Best: Claude Opus 5.5 · 1822
Multimodal
ChartographyChart tasks — The same chart-centred test with no tools: the AI has to read each chart unaided. Scores run far lower than the with-tools version, so read the two as separate tests. Higher is better.
15.6%no tools
#3 of 3Best: Claude Opus 5.5 · 64.4%
Robustness
BullshitBench v2Nonsense detection — Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
80%
#7 of 82Best: Claude Opus 4.8 · 95%
Gray Swan IPIPrompt injection robustness — Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. Gray Swan's indirect prompt injection benchmark measures how often such an attack succeeds when the attacker gets a single try. Lower is better.
0.6%k = 1
4.7%k = 10
5.9%k = 15
k = 1
#4 of 13Best: Claude Opus 5 · 0.2%
k = 10
#4 of 13Best: Claude Opus 5 · 1.6%
k = 15
#4 of 13Best: Claude Opus 5 · 2%
Community preference
threejsevalCommunity preference (Three.js) — Every model gets the same prompt — "the Eiffel Tower", "a glass fishbowl", "a robot arm picking toys into a box" — and builds a 3D scene in Three.js. Real people then see two scenes side by side, names hidden, and vote for the one they prefer. The votes become a chess-style Elo rating on threejseval.com, averaged across all the prompts. It measures whether the scene looks and moves right to a human eye, not whether the code passes a test. Higher is better.
1246
#17 of 18Best: Claude Opus 5.5 · 2075
Source: CursorBench, retrieved 28 September 2026 · Source: Terminal-Bench, retrieved 30 August 2026 · Source: BenchLM, retrieved 24 August 2026. Other sources are identified on the linked benchmark pages.
Compare Claude Sonnet 5 with
Suggested comparisons
Claude Sonnet 5vsClaude Sonnet 5.5Claude Sonnet 5vsGPT-6 SolClaude Sonnet 5vsGemini 3.8 FlashClaude Sonnet 5vsMuse Spark 1.3Claude Sonnet 5vsGrok 4.7Claude Sonnet 5vsDeepSeek-V4.1-FlashClaude Sonnet 5vsMistral Medium 3.5Claude Sonnet 5vsKimi K3Claude Sonnet 5vsGLM-5.3-FlashClaude Sonnet 5vsQwen3.8-Max-0902Claude Sonnet 5vsNemotron 3.5 LightningFrequently asked questions
Claude Sonnet 5 was released by Anthropic on Tuesday, Jun 30 2026.
All Anthropic releases
29 tracked2026
12 releasesClaude Sonnet 5.5
Sep 28 2026
Claude Opus 5.5
Sep 22 2026
Claude Fable 5.1
Sep 1 2026
Claude Mythos 5.1
Sep 1 2026
Claude Opus 5
Jul 24 2026
Claude Sonnet 5
Jun 30 2026
Claude Fable 5
Jun 9 2026
Claude Mythos 5
Jun 9 2026
Claude Opus 4.8
May 28 2026
Claude Opus 4.7
Apr 16 2026
Claude Sonnet 4.6
Feb 17 2026
Claude Opus 4.6
Feb 5 2026