Claude Opus 4.6vsGrok 4.20 Beta

Claude Opus 4.6
Grok 4.20 Beta
Benchmarks
Nonsense detection
BullshitBench v2
87%Best
56%
Coding
SWE-Bench Verified
80.8%
Next.js coding
Next.js Evals
75%
Web browsing
BrowseComp
83.7%
Multidisciplinary reasoning
Humanity's Last Exam · with tools
53%
Abstract reasoning
ARC-AGI-2
53.3%
Science
GPQA Diamond
91.3%
Community preference
Arena Elo (Text)
1504Best
1475
Community preference (code)
Arena Elo (Code)
1543
Overview
CompanyAnthropicSpaceXAI
Release dateFeb 5 2026Feb 17 2026
AccessProprietaryProprietary

Which is better: Claude Opus 4.6 or Grok 4.20 Beta?

Claude Opus 4.6 leads Grok 4.20 Beta on 2 of the 2 benchmarks they both report (BullshitBench v2, Arena Elo (Text)). Claude Opus 4.6 shipped 12 days before Grok 4.20 Beta, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude Opus 4.6 leads at 87% vs Grok 4.20 Beta at 56%. On Arena Elo (Text), Claude Opus 4.6 leads at 1504 vs Grok 4.20 Beta at 1475.

Frequently asked questions

Claude Opus 4.6 was released by Anthropic on Feb 5 2026.

Grok 4.20 Beta was released by SpaceXAI on Feb 17 2026.

Other comparisons