Gemini 3.1 ProvsGrok 4.20 Beta

Gemini 3.1 Pro
Grok 4.20 Beta
Benchmarks
Nonsense detection
BullshitBench v2
37%
56%Best
Prompt injection robustness
Gray Swan IPI · k = 1
14.2%
Prompt injection robustness
Gray Swan IPI · k = 10
45.7%
Prompt injection robustness
Gray Swan IPI · k = 15
49.2%
Agentic coding
SWE-Bench Pro
54.2%
Coding
SWE-Bench Verified
80.6%
Agentic coding
DeepSWE 1.1
12%
ML engineering
MLE-Bench
42.6%
Next.js coding
Next.js Evals
75%
Agentic terminal coding
Terminal-Bench 2.1
70.3%
Agentic terminal coding
Terminal-Bench 2.0
68.5%
Multi-step tool use
MCP Atlas
78.2%
General tool use
Toolathlon
48.8%
Web browsing
BrowseComp
85.9%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
44.4%
Multidisciplinary reasoning
Humanity's Last Exam · with tools
51.4%
Abstract reasoning
ARC-AGI-2
77.1%Best
53.3%
Advanced math
FrontierMath · Tier 1–3
36.9%
Advanced math
FrontierMath · Tier 4
16.7%
Science
GPQA Diamond
94.3%
Agentic computer use
OSWorld-Verified
76.2%
Agentic financial analysis
Finance Agent v2
43%
Knowledge work
GDPval-AA
1314
Knowledge work
GDPval-AA v2
965
Knowledge work
GDPval (win/tie rate)
67.3%
Chart reasoning
CharXiv Reasoning
83.3%
Multimodal reasoning
MMMU-Pro
80.5%
Spatial reasoning
Blueprint-Bench 2
26.5%
Long context
MRCR v2 (8-needle) · 128k average
84.9%
Long context
MRCR v2 (8-needle) · 1M pointwise
26.3%
Community preference
Arena Elo (Text)
1485Best
1475
Community preference (code)
Arena Elo (Code)
1445
Overview
CompanyGoogleSpaceXAI
Release dateFeb 19 2026Feb 17 2026
AccessProprietaryProprietary

Which is better: Gemini 3.1 Pro or Grok 4.20 Beta?

Gemini 3.1 Pro leads Grok 4.20 Beta on 2 of the 3 benchmarks they both report (BullshitBench v2, ARC-AGI-2, Arena Elo (Text)). Grok 4.20 Beta shipped 2 days before Gemini 3.1 Pro, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Grok 4.20 Beta leads at 56% vs Gemini 3.1 Pro at 37%. On ARC-AGI-2, Gemini 3.1 Pro leads at 77.1% vs Grok 4.20 Beta at 53.3%. On Arena Elo (Text), Gemini 3.1 Pro leads at 1485 vs Grok 4.20 Beta at 1475.

Frequently asked questions

Gemini 3.1 Pro was released by Google on Feb 19 2026.

Grok 4.20 Beta was released by SpaceXAI on Feb 17 2026.

Gemini 3.1 Pro leads on ARC-AGI-2 — Gemini 3.1 Pro 77.1% vs Grok 4.20 Beta 53.3%.

Other comparisons