Claude Opus 4.8vsGrok 4.20 Beta

Claude Opus 4.8
Grok 4.20 Beta
Specifications
Context window
1M
API pricing
Input price
$5.00$1.25
Output price
$25.00$2.50
Cached input price
$0.50$0.20
Benchmarks
BullshitBench v2
95%56%
Arena Elo (Text)
14821475
Benchmarks
Gray Swan IPI · k = 1
0.5%
Gray Swan IPI · k = 10
4.1%
Gray Swan IPI · k = 15
5.5%
SWE-Bench Pro
69.2%
SWE-Bench Multilingual
84.4%
CursorBench v3.2
62.3%
CursorBench v3.1
63.8%
DeepSWE 1.0
55.8%
Next.js Evals
88%
Frontier-Bench v0.1
21.1%
Terminal-Bench 2.1
74.6%
BU Bench
74%
BrowseComp
84.3%
Humanity's Last Exam · no tools
49.8%
Humanity's Last Exam · with tools
57.9%
ARC-AGI-2
53.3%
OSWorld-Verified
83.4%
Finance Agent v2
53.9%
Harvey's Legal Agent Benchmark
9.58%
TaxEval v2
75.63%
MedScribe
85.75%
GDPval-AA
1890
GDPval-AA v2
1600
Arena Elo (Code)
1568
Overview
CompanyAnthropicSpaceXAI
Release dateMay 28 2026Feb 17 2026
AccessProprietaryProprietary

Frequently asked questions

Claude Opus 4.8 leads Grok 4.20 Beta on 2 of the 2 benchmarks they both report (BullshitBench v2, Arena Elo (Text)). Grok 4.20 Beta is cheaper on both input and output: $1.25 vs $5.00 per million input tokens, and $2.50 vs $25.00 per million output tokens. Figures are base-tier rates. Grok 4.20 Beta shipped 100 days before Claude Opus 4.8, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude Opus 4.8 leads at 95% vs Grok 4.20 Beta at 56%. On Arena Elo (Text), Claude Opus 4.8 leads at 1482 vs Grok 4.20 Beta at 1475.

Claude Opus 4.8 was released by Anthropic on May 28 2026.

Grok 4.20 Beta was released by SpaceXAI on Feb 17 2026.

Grok 4.20 Beta is cheaper on both input and output: $1.25 vs $5.00 per million input tokens, and $2.50 vs $25.00 per million output tokens. Figures are base-tier rates. Rates are pay-as-you-go API prices verified on August 18, 2026.

Other comparisons