Claude 3.7 SonnetvsClaude Opus 4.8

Claude 3.7 Sonnet
Claude Opus 4.8
Specifications
Context window
1M
Benchmarks
Nonsense detection
BullshitBench v2
49%
95%Best
Prompt injection robustness
Gray Swan IPI · k = 1
0.5%
Prompt injection robustness
Gray Swan IPI · k = 10
4.1%
Prompt injection robustness
Gray Swan IPI · k = 15
5.5%
Agentic coding
SWE-Bench Pro
69.2%
Coding
SWE-Bench Verified
62.3%
Multilingual coding
SWE-Bench Multilingual
84.4%
Agentic coding
CursorBench v3.2
62.3%
Agentic coding
CursorBench v3.1
63.8%
Agentic coding
DeepSWE 1.0
55.8%
Next.js coding
Next.js Evals
88%
Agentic computer work
Frontier-Bench v0.1
21.1%
Agentic terminal coding
Terminal-Bench 2.1
74.6%
Browser agent
BU Bench
74%
Web browsing
BrowseComp
84.3%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
49.8%
Multidisciplinary reasoning
Humanity's Last Exam · with tools
57.9%
Science
GPQA Diamond
68%
Agentic computer use
OSWorld-Verified
83.4%
Agentic financial analysis
Finance Agent v2
53.9%
Agentic legal work
Harvey's Legal Agent Benchmark
9.58%
Tax questions
TaxEval v2
75.63%
Medical admin work
MedScribe
85.75%
Knowledge work
GDPval-AA
1890
Knowledge work
GDPval-AA v2
1600
Community preference
Arena Elo (Text)
1482
Community preference (code)
Arena Elo (Code)
1568
Overview
CompanyAnthropicAnthropic
Release dateFeb 24 2025May 28 2026
AccessProprietaryProprietary

Which is better: Claude 3.7 Sonnet or Claude Opus 4.8?

Claude Opus 4.8 leads Claude 3.7 Sonnet on 1 of the 1 benchmark they both report (BullshitBench v2). Claude 3.7 Sonnet shipped 458 days before Claude Opus 4.8, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude Opus 4.8 leads at 95% vs Claude 3.7 Sonnet at 49%.

Frequently asked questions

Claude 3.7 Sonnet was released by Anthropic on Feb 24 2025.

Claude Opus 4.8 was released by Anthropic on May 28 2026.

Other comparisons