Claude Sonnet 5
Claude Sonnet 5 is an AI model released by Anthropic on Tuesday, Jun 30 2026, 21 days after Claude Fable 5. Benchmark results (shown below) cover Supabase Evals, BullshitBench v2, Gray Swan IPI, SWE-Bench Pro, CursorBench v3.2, CursorBench v3.1, and 9 more.
Benchmarks
Supabase coding
Supabase EvalsSupabase's own open benchmark: a coding agent is dropped into a real Supabase project and asked to do real work — set up a schema, fix a broken security policy, debug an Edge Function — and every run is checked against a live Supabase stack. This is the headline number, where the agent has Supabase's own skills loaded, as most people building on Supabase would. The score is the share of scenarios it got right. Higher is better.
95.5%
with skills
90.9%
no skills
with skills
#1Best published Supabase score of all tracked models
no skills
#1Best published Supabase score of all tracked models
Nonsense detection
BullshitBench v2Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
80%
#6 of 67Best: Claude Opus 4.8 · 95%
Prompt injection robustness
Gray Swan IPIAttackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. Gray Swan's indirect prompt injection benchmark measures how often such an attack succeeds when the attacker gets a single try. Lower is better.
0.6%
k = 1
4.7%
k = 10
5.9%
k = 15
k = 1
#4 of 13Best: Claude Opus 5 · 0.2%
k = 10
#4 of 13Best: Claude Opus 5 · 1.6%
k = 15
#4 of 13Best: Claude Opus 5 · 2%
Agentic coding
SWE-Bench ProCan the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better.
63.2%
#6 of 20Best: Claude Fable 5 · 80.3%
Agentic coding
CursorBench v3.2Cursor's own test of harder, real-world coding tasks inside a code editor, on the refreshed v3.2 task set. Scores aren't comparable with v3.1. Higher is better.
61.5%
#8 of 15Best: Claude Fable 5 · 70.5%
Agentic coding
CursorBench v3.1Cursor's own test of harder, real-world coding tasks inside a code editor. Higher is better.
61.2%
#6 of 12Best: Claude Fable 5 · 72.9%
Next.js coding
Next.js EvalsVercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better.
79%
#12 of 24Best: Composer 2.5 · 92%
Agentic computer work
Frontier-Bench v0.1A hard, ever-evolving set of real computer tasks — coding, system administration, data work, and more — that an AI agent has to complete on its own. Run by the Harbor / Laude Institute team as the successor to Terminal-Bench (v0.1 is the first release of the task set). The score is the share of tasks solved. Higher is better.
14.6%
#7 of 9Best: Claude Opus 5 · 43.3%
Agentic terminal coding
Terminal-Bench 2.1Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better.
80.4%
#13 of 25Best: GPT-5.6 Sol · 88.8%
Web browsing
BrowseCompCan the AI browse the web and track down hard-to-find answers? Higher is better.
84.7%
#7 of 17Best: Kimi K3 · 91.2%
Multidisciplinary reasoning
Humanity's Last ExamHumanity's Last Exam — extremely hard expert questions across many subjects, written so you can't just look up the answer. “No tools” means the AI answers on its own. Higher is better.
43.2%
no tools
57.4%
with tools
no tools
#6 of 19Best: Claude Opus 5 · 56.3%
with tools
#8 of 21Best: Claude Opus 5 · 64.7%
Agentic computer use
OSWorld-VerifiedCan the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Higher is better.
81.2%
#5 of 17Best: Qwen3.8-Max · 86.1%
Knowledge work
GDPval-AAMeasures how well the AI does economically valuable knowledge work, judged against human experts. Shown as a rating (like a chess Elo) — higher is better.
1618
#7 of 9Best: Claude Fable 5 · 1932
Community preference
Arena Elo (Text)Real people chat with two anonymous AIs side by side and vote for the answer they prefer. Votes become a chess-style Elo rating on arena.ai — it measures which AI people actually like, not test scores. Higher is better.
1463
#27 of 31Best: Claude Fable 5 · 1509
Community preference (code)
Arena Elo (Code)Like the text arena, but people vote on which AI writes better code. The votes become a chess-style Elo rating on arena.ai. Higher is better.
1543
#13 of 48Best: Kimi K3 · 1679
Claude Sonnet 5 — frequently asked questions
- When was Claude Sonnet 5 released?
- Claude Sonnet 5 was released by Anthropic on Tuesday, Jun 30 2026.
- Who made Claude Sonnet 5?
- Claude Sonnet 5 was built by Anthropic. AI safety company building the Claude family of models. Founded in 2021 by former OpenAI researchers.
- What benchmark scores did Claude Sonnet 5 get?
- Claude Sonnet 5 reports 19 tracked benchmark scores — BullshitBench v2: 80%; Gray Swan IPI (k = 1): 0.6%; Gray Swan IPI (k = 10): 4.7%; Gray Swan IPI (k = 15): 5.9%; SWE-Bench Pro: 63.2%; CursorBench v3.2: 61.5%; CursorBench v3.1: 61.2%; Next.js Evals: 79%; Supabase Evals (with skills): 95.5%; Supabase Evals (no skills): 90.9%; Frontier-Bench v0.1: 14.6%; Terminal-Bench 2.1: 80.4%; BrowseComp: 84.7%; Humanity's Last Exam (no tools): 43.2%; Humanity's Last Exam (with tools): 57.4%; OSWorld-Verified: 81.2%; GDPval-AA: 1618; Arena Elo (Text): 1463; Arena Elo (Code): 1543. Scores are the figures published at release by Anthropic.
- Is Claude Sonnet 5 open source?
- No. Claude Sonnet 5 is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.
- What came before and after Claude Sonnet 5?
- Anthropic's previous tracked release was Claude Fable 5 on Jun 9 2026, 21 days earlier. It was followed by Claude Opus 5 on Jul 24 2026.
Compare Claude Sonnet 5 with
Claude Sonnet 5vsClaude Opus 5Claude Sonnet 5vsGPT-5.6-CyberClaude Sonnet 5vsGemini 3.7 FlashClaude Sonnet 5vsMuse GlimmerClaude Sonnet 5vsGrok 4.6Claude Sonnet 5vsDeepSeek-V4-Pro-0813Claude Sonnet 5vsMistral Medium 3.5Claude Sonnet 5vsKimi K3Claude Sonnet 5vsComposer 2.5Claude Sonnet 5vsGLM-5.3Claude Sonnet 5vsQwen3.8-27B