Claude Sonnet 4.6vsGemini 3.7 Flash

Claude Sonnet 4.6
Gemini 3.7 Flash
Specifications
Context window
1M
API pricing
Input price
$3.00$0.75
Output price
$15.00$3.75
Cached input price
$0.30$0.075
Cheapest input
$3.00Amazon Bedrock
Cheapest output
$15.00Amazon Bedrock
Benchmarks
BullshitBench v2
91%35%
ProgramBench
0%0%
DeepSWE 1.1
30%65.3%
CharXiv Reasoning
72.4%84.5%
MRCR v2 (8-needle) · 128k average
84.9%97%
Benchmarks
SWE-Bench Verified
79.6%
FrontierCode v1.1 (Main) · main split
43.6%
Next.js Evals
45%
Terminal-Bench 4.0
11.21%
Terminal-Bench 3.0
14.9%
Terminal-Bench 2.1
85.8%
MCP Atlas
69.5%
BU Bench
62%
Humanity's Last Exam · no tools
33.2%
Humanity's Last Exam · with tools
49%
Humanity's Last Exam (Verified)
53.6%
ARC-AGI-2
58.3%
BioMysteryBench · hard
43.5%
BioMysteryBench · human solved
87.1%
LAB-Bench 2
82.1%
GPQA Diamond
89.9%
OSWorld 2.0
38.1%
OSWorld-Verified
72.5%
Agent's Last Exam · pass@1
26.3%
AutomationBench
30.4%
Finance Agent v2
51%
Harvey's Legal Agent Benchmark
8.8%
AA Intelligence Index
56
GDPval-AA
1676
GDPval-AA v2
1525
GDP.PDF
34%
LVBench
85.4%
MMMU-Pro
74.5%
Blueprint-Bench 2
6.7%
MRCR v2 (8-needle) · 1M pointwise
62.5%
threejseval
1526
Overview
CompanyAnthropicGoogle
Release dateFeb 17 2026Aug 13 2026
AccessProprietaryProprietary

Other comparisons

Claude Sonnet 4.6vsGPT-6 AstraGemini 3.7 FlashvsGPT-6 AstraClaude Sonnet 4.6vsMuse Spark 1.3Gemini 3.7 FlashvsMuse Spark 1.3Claude Sonnet 4.6vsGrok 4.6Gemini 3.7 FlashvsGrok 4.6Claude Sonnet 4.6vsDeepSeek-V4.1-FlashGemini 3.7 FlashvsDeepSeek-V4.1-FlashClaude Sonnet 4.6vsMistral Medium 3.5Gemini 3.7 FlashvsMistral Medium 3.5Claude Sonnet 4.6vsKimi K3Gemini 3.7 FlashvsKimi K3

Frequently asked questions

Gemini 3.7 Flash leads Claude Sonnet 4.6 on 3 of the 5 benchmarks they both report. Gemini 3.7 Flash is cheaper on both input and output: $0.75 vs $3.00 per million input tokens, and $3.75 vs $15.00 per million output tokens. Claude Sonnet 4.6 shipped 177 days before Gemini 3.7 Flash, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude Sonnet 4.6 leads at 91% vs Gemini 3.7 Flash at 35%. On ProgramBench, both models score 0%. On DeepSWE 1.1, Gemini 3.7 Flash leads at 65.3% vs Claude Sonnet 4.6 at 30%. On CharXiv Reasoning, Gemini 3.7 Flash leads at 84.5% vs Claude Sonnet 4.6 at 72.4%. On MRCR v2 (8-needle) · 128k average, Gemini 3.7 Flash leads at 97% vs Claude Sonnet 4.6 at 84.9%.