Gemini 3.7 FlashvsGrok 4.5

Gemini 3.7 Flash
Grok 4.5
Specifications
Context window
1M
Benchmarks
Nonsense detection
BullshitBench v2
54%
Prompt injection robustness
Gray Swan IPI · k = 1
13.4%
Prompt injection robustness
Gray Swan IPI · k = 10
54.2%
Prompt injection robustness
Gray Swan IPI · k = 15
60.8%
Agentic coding
SWE-Bench Pro
64.7%
Multilingual coding
SWE-Bench Multilingual
78%
Agentic coding
CursorBench v3.2
66.7%
Agentic coding
DeepSWE 1.1
65.3%Best
54%
Agentic coding
DeepSWE 1.0
62%
Agentic coding
FrontierCode v1.1 (Main) · main split
43.6%
Agentic coding
FrontierCode v1.1 (Extended) · extended split
56.6%
Expert software engineering
APEX-SWE
53.6%
Next.js coding
Next.js Evals
83%
Agentic computer work
Frontier-Bench v0.1
17.8%
Agentic terminal coding
Terminal-Bench 3.0
14.9%
15.7%Best
Agentic terminal coding
Terminal-Bench 2.1
85.8%Best
83.3%
Expert agentic work
APEX-Agents
47.1%
Multidisciplinary reasoning
Humanity's Last Exam (Verified)
53.6%
Biology
BioMysteryBench · hard
43.5%
Biology
BioMysteryBench · human solved
87.1%
Biology
LAB-Bench 2
82.1%
Agentic computer use
OSWorld 2.0
38.1%
Agentic computer use
Agent's Last Exam
26.3%
Business workflows
AutomationBench
30.4%
Agentic legal work
Harvey's Legal Agent Benchmark
90.7%Best
12.92%
Medical admin work
MedScribe
86.88%
Overall intelligence
AA Intelligence Index
56
56
Knowledge work
GDPval-AA v2
1525
1526Best
Knowledge work
AA-Briefcase
1313
Chart reasoning
CharXiv Reasoning
84.5%
Document comprehension
GDP.PDF
34%
Video understanding
LVBench
85.4%
Long context
MRCR v2 (8-needle) · 128k average
97%
Long context
MRCR v2 (8-needle) · 1M pointwise
62.5%
Community preference
Arena Elo (Text)
1468
Community preference (code)
Arena Elo (Code)
1588Best
1549
Overview
CompanyGoogleSpaceXAI
Release dateAug 13 2026Jul 8 2026
AccessProprietaryProprietary

Which is better: Gemini 3.7 Flash or Grok 4.5?

Gemini 3.7 Flash leads Grok 4.5 on 4 of the 7 benchmarks they both report. Grok 4.5 shipped 36 days before Gemini 3.7 Flash, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On DeepSWE 1.1, Gemini 3.7 Flash leads at 65.3% vs Grok 4.5 at 54%. On Terminal-Bench 3.0, Grok 4.5 leads at 15.7% vs Gemini 3.7 Flash at 14.9%. On Terminal-Bench 2.1, Gemini 3.7 Flash leads at 85.8% vs Grok 4.5 at 83.3%. On Harvey's Legal Agent Benchmark, Gemini 3.7 Flash leads at 90.7% vs Grok 4.5 at 12.92%. On AA Intelligence Index, both models score 56. On GDPval-AA v2, Grok 4.5 leads at 1526 vs Gemini 3.7 Flash at 1525. On Arena Elo (Code), Gemini 3.7 Flash leads at 1588 vs Grok 4.5 at 1549.

Frequently asked questions

Gemini 3.7 Flash was released by Google on Aug 13 2026.

Grok 4.5 was released by SpaceXAI on Jul 8 2026.

Gemini 3.7 Flash leads on DeepSWE 1.1 — Gemini 3.7 Flash 65.3% vs Grok 4.5 54%.

Other comparisons