Anthropic
Claude Sonnet 4
Claude Sonnet 4 is an AI model released by Anthropic on Thursday, May 22 2025, 87 days after Claude 3.7 Sonnet. Benchmark results (shown below) cover BullshitBench v2, SWE-Bench Verified, and GPQA Diamond.
Benchmarks
Nonsense detection
BullshitBench v2Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
30%
#32 of 60Best: Claude Opus 4.8 · 95%
Coding
SWE-Bench VerifiedReal coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better.
72.7%
#20 of 33Best: Claude Fable 5 · 95.5%
Science
GPQA DiamondGraduate-level science questions in biology, physics, and chemistry — hard enough that subject-matter PhDs score around 65%. Higher is better.
75.4%
#28 of 37Best: GPT-5.4-Pro · 94.4%
Compare Claude Sonnet 4 with
Claude Sonnet 4 vs Claude Sonnet 5Claude Sonnet 4 vs GPT-5.6 SolClaude Sonnet 4 vs Gemini 3.5 FlashClaude Sonnet 4 vs Muse Spark 1.1Claude Sonnet 4 vs Grok 4.5Claude Sonnet 4 vs DeepSeek-V4-ProClaude Sonnet 4 vs Mistral Medium 3.5Claude Sonnet 4 vs Kimi K3Claude Sonnet 4 vs Composer 2.5Claude Sonnet 4 vs GLM-5.2