Claude 3.7 SonnetvsGemini 3.1 Pro

Claude 3.7 Sonnet
Gemini 3.1 Pro
Benchmarks
Nonsense detection
BullshitBench v2
49%Best
37%
Prompt injection robustness
Gray Swan IPI · k = 1
14.2%
Prompt injection robustness
Gray Swan IPI · k = 10
45.7%
Prompt injection robustness
Gray Swan IPI · k = 15
49.2%
Agentic coding
SWE-Bench Pro
54.2%
Coding
SWE-Bench Verified
62.3%
80.6%Best
Agentic coding
DeepSWE 1.1
12%
ML engineering
MLE-Bench
42.6%
Next.js coding
Next.js Evals
75%
Agentic terminal coding
Terminal-Bench 2.1
70.3%
Agentic terminal coding
Terminal-Bench 2.0
68.5%
Multi-step tool use
MCP Atlas
78.2%
General tool use
Toolathlon
48.8%
Web browsing
BrowseComp
85.9%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
44.4%
Multidisciplinary reasoning
Humanity's Last Exam · with tools
51.4%
Abstract reasoning
ARC-AGI-2
77.1%
Advanced math
FrontierMath · Tier 1–3
36.9%
Advanced math
FrontierMath · Tier 4
16.7%
Science
GPQA Diamond
68%
94.3%Best
Agentic computer use
OSWorld-Verified
76.2%
Agentic financial analysis
Finance Agent v2
43%
Knowledge work
GDPval-AA
1314
Knowledge work
GDPval-AA v2
965
Knowledge work
GDPval (win/tie rate)
67.3%
Chart reasoning
CharXiv Reasoning
83.3%
Multimodal reasoning
MMMU-Pro
80.5%
Spatial reasoning
Blueprint-Bench 2
26.5%
Long context
MRCR v2 (8-needle) · 128k average
84.9%
Long context
MRCR v2 (8-needle) · 1M pointwise
26.3%
Community preference
Arena Elo (Text)
1485
Community preference (code)
Arena Elo (Code)
1445
Overview
CompanyAnthropicGoogle
Release dateFeb 24 2025Feb 19 2026
AccessProprietaryProprietary

Which is better: Claude 3.7 Sonnet or Gemini 3.1 Pro?

Gemini 3.1 Pro leads Claude 3.7 Sonnet on 2 of the 3 benchmarks they both report (BullshitBench v2, SWE-Bench Verified, GPQA Diamond). Claude 3.7 Sonnet shipped 360 days before Gemini 3.1 Pro, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude 3.7 Sonnet leads at 49% vs Gemini 3.1 Pro at 37%. On SWE-Bench Verified, Gemini 3.1 Pro leads at 80.6% vs Claude 3.7 Sonnet at 62.3%. On GPQA Diamond, Gemini 3.1 Pro leads at 94.3% vs Claude 3.7 Sonnet at 68%.

Frequently asked questions

Claude 3.7 Sonnet was released by Anthropic on Feb 24 2025.

Gemini 3.1 Pro was released by Google on Feb 19 2026.

Gemini 3.1 Pro leads on SWE-Bench Verified — Claude 3.7 Sonnet 62.3% vs Gemini 3.1 Pro 80.6%.

Gemini 3.1 Pro leads on GPQA Diamond — Claude 3.7 Sonnet 68% vs Gemini 3.1 Pro 94.3%.

Other comparisons