Claude Opus 4.8vsGemini 3.7 Flash

Claude Opus 4.8
Gemini 3.7 Flash
Specifications
Context window
1M
1M
Benchmarks
Nonsense detection
BullshitBench v2
95%
Prompt injection robustness
Gray Swan IPI · k = 1
0.5%
Prompt injection robustness
Gray Swan IPI · k = 10
4.1%
Prompt injection robustness
Gray Swan IPI · k = 15
5.5%
Agentic coding
SWE-Bench Pro
69.2%
Multilingual coding
SWE-Bench Multilingual
84.4%
Agentic coding
CursorBench v3.2
62.3%
Agentic coding
CursorBench v3.1
63.8%
Agentic coding
DeepSWE 1.1
65.3%
Agentic coding
DeepSWE 1.0
55.8%
Agentic coding
FrontierCode v1.1 (Main) · main split
43.6%
Next.js coding
Next.js Evals
88%
Agentic computer work
Frontier-Bench v0.1
21.1%
Agentic terminal coding
Terminal-Bench 3.0
14.9%
Agentic terminal coding
Terminal-Bench 2.1
74.6%
85.8%Best
Browser agent
BU Bench
74%
Web browsing
BrowseComp
84.3%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
49.8%
Multidisciplinary reasoning
Humanity's Last Exam · with tools
57.9%
Multidisciplinary reasoning
Humanity's Last Exam (Verified)
53.6%
Biology
BioMysteryBench · hard
43.5%
Biology
BioMysteryBench · human solved
87.1%
Biology
LAB-Bench 2
82.1%
Agentic computer use
OSWorld 2.0
38.1%
Agentic computer use
OSWorld-Verified
83.4%
Agentic computer use
Agent's Last Exam
26.3%
Business workflows
AutomationBench
30.4%
Agentic financial analysis
Finance Agent v2
53.9%
Agentic legal work
Harvey's Legal Agent Benchmark
9.58%
90.7%Best
Tax questions
TaxEval v2
75.63%
Medical admin work
MedScribe
85.75%
Overall intelligence
AA Intelligence Index
56
Knowledge work
GDPval-AA
1890
Knowledge work
GDPval-AA v2
1600Best
1525
Chart reasoning
CharXiv Reasoning
84.5%
Document comprehension
GDP.PDF
34%
Video understanding
LVBench
85.4%
Long context
MRCR v2 (8-needle) · 128k average
97%
Long context
MRCR v2 (8-needle) · 1M pointwise
62.5%
Community preference
Arena Elo (Text)
1482
Community preference (code)
Arena Elo (Code)
1568
1588Best
Overview
CompanyAnthropicGoogle
Release dateMay 28 2026Aug 13 2026
AccessProprietaryProprietary

Which is better: Claude Opus 4.8 or Gemini 3.7 Flash?

Gemini 3.7 Flash leads Claude Opus 4.8 on 3 of the 4 benchmarks they both report (Terminal-Bench 2.1, Harvey's Legal Agent Benchmark, GDPval-AA v2, Arena Elo (Code)). Claude Opus 4.8 shipped 77 days before Gemini 3.7 Flash, so benchmark comparisons should account for the intervening progress.

Context windows are 1M (Claude Opus 4.8) vs 1M (Gemini 3.7 Flash).

On Terminal-Bench 2.1, Gemini 3.7 Flash leads at 85.8% vs Claude Opus 4.8 at 74.6%. On Harvey's Legal Agent Benchmark, Gemini 3.7 Flash leads at 90.7% vs Claude Opus 4.8 at 9.58%. On GDPval-AA v2, Claude Opus 4.8 leads at 1600 vs Gemini 3.7 Flash at 1525. On Arena Elo (Code), Gemini 3.7 Flash leads at 1588 vs Claude Opus 4.8 at 1568.

Frequently asked questions

Claude Opus 4.8 was released by Anthropic on May 28 2026.

Gemini 3.7 Flash was released by Google on Aug 13 2026.

Gemini 3.7 Flash leads on Terminal-Bench 2.1 — Claude Opus 4.8 74.6% vs Gemini 3.7 Flash 85.8%.

Claude Opus 4.8 has a 1M context window; Gemini 3.7 Flash has 1M.

Other comparisons