Claude Sonnet 4.6vsGrok 4.20 Beta

Claude Sonnet 4.6
Grok 4.20 Beta
Benchmarks
Nonsense detection
BullshitBench v2
91%Best
56%
Coding
SWE-Bench Verified
79.6%
Agentic coding
CursorBench v3.1
49%
Agentic coding
DeepSWE 1.1
30%
Next.js coding
Next.js Evals
58%
Multi-step tool use
MCP Atlas
69.5%
Browser agent
BU Bench
62%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
33.2%
Abstract reasoning
ARC-AGI-2
58.3%Best
53.3%
Science
GPQA Diamond
89.9%
Agentic computer use
OSWorld-Verified
72.5%
Agentic financial analysis
Finance Agent v2
51%
Knowledge work
GDPval-AA
1676
Chart reasoning
CharXiv Reasoning
72.4%
Multimodal reasoning
MMMU-Pro
74.5%
Spatial reasoning
Blueprint-Bench 2
6.7%
Long context
MRCR v2 (8-needle) · 128k average
84.9%
Community preference
Arena Elo (Text)
1475
Community preference (code)
Arena Elo (Code)
1521
Overview
CompanyAnthropicSpaceXAI
Release dateFeb 17 2026Feb 17 2026
AccessProprietaryProprietary

Which is better: Claude Sonnet 4.6 or Grok 4.20 Beta?

Claude Sonnet 4.6 leads Grok 4.20 Beta on 2 of the 2 benchmarks they both report (BullshitBench v2, ARC-AGI-2). Both models were released on the same day — Feb 17 2026.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude Sonnet 4.6 leads at 91% vs Grok 4.20 Beta at 56%. On ARC-AGI-2, Claude Sonnet 4.6 leads at 58.3% vs Grok 4.20 Beta at 53.3%.

Frequently asked questions

Claude Sonnet 4.6 was released by Anthropic on Feb 17 2026.

Grok 4.20 Beta was released by SpaceXAI on Feb 17 2026.

Claude Sonnet 4.6 leads on ARC-AGI-2 — Claude Sonnet 4.6 58.3% vs Grok 4.20 Beta 53.3%.

Other comparisons