Claude Sonnet 5vso3

Claude Sonnet 5
o3
Benchmarks
Nonsense detection
BullshitBench v2
80%Best
26%
Prompt injection robustness
Gray Swan IPI · k = 1
0.6%
Prompt injection robustness
Gray Swan IPI · k = 10
4.7%
Prompt injection robustness
Gray Swan IPI · k = 15
5.9%
Agentic coding
SWE-Bench Pro
63.2%
Agentic coding
CursorBench v3.2
61.5%
Agentic coding
CursorBench v3.1
61.2%
Next.js coding
Next.js Evals
79%
Supabase coding
Supabase Evals · with skills
95.5%
Supabase coding
Supabase Evals · no skills
90.9%
Agentic computer work
Frontier-Bench v0.1
14.6%
Agentic terminal coding
Terminal-Bench 2.1
80.4%
Web browsing
BrowseComp
84.7%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
43.2%
Multidisciplinary reasoning
Humanity's Last Exam · with tools
57.4%
Agentic computer use
OSWorld-Verified
81.2%
Knowledge work
GDPval-AA
1618
Community preference
Arena Elo (Text)
1463
Community preference (code)
Arena Elo (Code)
1543
Overview
CompanyAnthropicOpenAI
Release dateJun 30 2026Apr 16 2025
AccessProprietaryProprietary

Which is better: Claude Sonnet 5 or o3?

Claude Sonnet 5 leads o3 on 1 of the 1 benchmark they both report (BullshitBench v2). o3 shipped 440 days before Claude Sonnet 5, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude Sonnet 5 leads at 80% vs o3 at 26%.

Frequently asked questions

Claude Sonnet 5 was released by Anthropic on Jun 30 2026.

o3 was released by OpenAI on Apr 16 2025.

Other comparisons