Claude Opus 4.8vsGrok 4.6

Claude Opus 4.8
Grok 4.6
Specifications
Context window
1M
Benchmarks
Nonsense detection
BullshitBench v2
95%
Prompt injection robustness
Gray Swan IPI · k = 1
0.5%
Prompt injection robustness
Gray Swan IPI · k = 10
4.1%
Prompt injection robustness
Gray Swan IPI · k = 15
5.5%
Agentic coding
SWE-Bench Pro
69.2%
Multilingual coding
SWE-Bench Multilingual
84.4%
Agentic coding
CursorBench v3.2
62.3%
69.9%Best
Agentic coding
CursorBench v3.1
63.8%
Agentic coding
DeepSWE 1.1
65.9%
Agentic coding
DeepSWE 1.0
55.8%
Agentic coding
FrontierCode v1.1 (Extended) · extended split
61.3%
Expert software engineering
APEX-SWE
56.4%
Next.js coding
Next.js Evals
88%
Agentic computer work
Frontier-Bench v0.1
21.1%
Agentic terminal coding
Terminal-Bench 3.0
26%
Agentic terminal coding
Terminal-Bench 2.1
74.6%
Expert agentic work
APEX-Agents
57.5%
Browser agent
BU Bench
74%
Web browsing
BrowseComp
84.3%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
49.8%
Multidisciplinary reasoning
Humanity's Last Exam · with tools
57.9%
Agentic computer use
OSWorld-Verified
83.4%
Agentic financial analysis
Finance Agent v2
53.9%
Agentic legal work
Harvey's Legal Agent Benchmark
9.58%
15.8%Best
Tax questions
TaxEval v2
75.63%
Medical admin work
MedScribe
85.75%
Overall intelligence
AA Intelligence Index
61
Knowledge work
GDPval-AA
1890
Knowledge work
GDPval-AA v2
1600
1753Best
Knowledge work
AA-Briefcase
1577
Community preference
Arena Elo (Text)
1482
Community preference (code)
Arena Elo (Code)
1568
Overview
CompanyAnthropicSpaceXAI
Release dateMay 28 2026Aug 12 2026
AccessProprietaryProprietary

Which is better: Claude Opus 4.8 or Grok 4.6?

Grok 4.6 leads Claude Opus 4.8 on 3 of the 3 benchmarks they both report (CursorBench v3.2, Harvey's Legal Agent Benchmark, GDPval-AA v2). Claude Opus 4.8 shipped 76 days before Grok 4.6, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On CursorBench v3.2, Grok 4.6 leads at 69.9% vs Claude Opus 4.8 at 62.3%. On Harvey's Legal Agent Benchmark, Grok 4.6 leads at 15.8% vs Claude Opus 4.8 at 9.58%. On GDPval-AA v2, Grok 4.6 leads at 1753 vs Claude Opus 4.8 at 1600.

Frequently asked questions

Claude Opus 4.8 was released by Anthropic on May 28 2026.

Grok 4.6 was released by SpaceXAI on Aug 12 2026.

Grok 4.6 leads on CursorBench v3.2 — Claude Opus 4.8 62.3% vs Grok 4.6 69.9%.

Other comparisons