Compare AI models

Specifications
Context window
—1M
API pricing
Input price
$3.00$0.75
Output price
$15.00$3.75
Cached input price
$0.30$0.075
Cheapest input
$3.00Amazon Bedrock$0.375Google
Cheapest output
$15.00Amazon Bedrock$1.875Google
Benchmarks
BullshitBench v2
91%35%
ProgramBench
0%0%
DeepSWE 1.1
30%65.3%
ARC-AGI-2
58.3%84.6%
CharXiv Reasoning
72.4%84.5%
MRCR v2 (8-needle) · 128k average
84.9%97%
Overview
CompanyAnthropicGoogle
Release dateFeb 17 2026Aug 13 2026
AccessClosedClosed
Model detailsView modelView model

Frequently asked questions

Gemini 3.7 Flash leads Claude Sonnet 4.6 on 4 of the 6 benchmarks they both report. Gemini 3.7 Flash is cheaper on both input and output: $0.75 vs $3.00 per million input tokens, and $3.75 vs $15.00 per million output tokens. Claude Sonnet 4.6 shipped 177 days before Gemini 3.7 Flash, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude Sonnet 4.6 leads at 91% vs Gemini 3.7 Flash at 35%. On ProgramBench, both models score 0%. On DeepSWE 1.1, Gemini 3.7 Flash leads at 65.3% vs Claude Sonnet 4.6 at 30%. On ARC-AGI-2, Gemini 3.7 Flash leads at 84.6% vs Claude Sonnet 4.6 at 58.3%. On CharXiv Reasoning, Gemini 3.7 Flash leads at 84.5% vs Claude Sonnet 4.6 at 72.4%. On MRCR v2 (8-needle) · 128k average, Gemini 3.7 Flash leads at 97% vs Claude Sonnet 4.6 at 84.9%.