Claude 3.7 SonnetvsClaude Sonnet 4.6

Claude 3.7 Sonnet
Claude Sonnet 4.6
Benchmarks
Nonsense detection
BullshitBench v2
49%
91%Best
Coding
SWE-Bench Verified
62.3%
79.6%Best
Agentic coding
CursorBench v3.1
49%
Agentic coding
DeepSWE 1.1
30%
Next.js coding
Next.js Evals
58%
Multi-step tool use
MCP Atlas
69.5%
Browser agent
BU Bench
62%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
33.2%
Abstract reasoning
ARC-AGI-2
58.3%
Science
GPQA Diamond
68%
89.9%Best
Agentic computer use
OSWorld-Verified
72.5%
Agentic financial analysis
Finance Agent v2
51%
Knowledge work
GDPval-AA
1676
Chart reasoning
CharXiv Reasoning
72.4%
Multimodal reasoning
MMMU-Pro
74.5%
Spatial reasoning
Blueprint-Bench 2
6.7%
Long context
MRCR v2 (8-needle) · 128k average
84.9%
Community preference
Arena Elo (Text)
1472
Community preference (code)
Arena Elo (Code)
1521
Overview
CompanyAnthropicAnthropic
Release dateFeb 24 2025Feb 17 2026
AccessProprietaryProprietary

Which is better: Claude 3.7 Sonnet or Claude Sonnet 4.6?

Claude Sonnet 4.6 leads Claude 3.7 Sonnet on 3 of the 3 benchmarks they both report (BullshitBench v2, SWE-Bench Verified, GPQA Diamond). Claude 3.7 Sonnet shipped 358 days before Claude Sonnet 4.6, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude Sonnet 4.6 leads at 91% vs Claude 3.7 Sonnet at 49%. On SWE-Bench Verified, Claude Sonnet 4.6 leads at 79.6% vs Claude 3.7 Sonnet at 62.3%. On GPQA Diamond, Claude Sonnet 4.6 leads at 89.9% vs Claude 3.7 Sonnet at 68%.

Frequently asked questions

Claude 3.7 Sonnet was released by Anthropic on Feb 24 2025.

Claude Sonnet 4.6 was released by Anthropic on Feb 17 2026.

Claude Sonnet 4.6 leads on SWE-Bench Verified — Claude 3.7 Sonnet 62.3% vs Claude Sonnet 4.6 79.6%.

Claude Sonnet 4.6 leads on GPQA Diamond — Claude 3.7 Sonnet 68% vs Claude Sonnet 4.6 89.9%.

Other comparisons