Claude Opus 4.8vsGrok 4.5

Claude Opus 4.8
Grok 4.5
Specifications
Context window
1M
Benchmarks
Nonsense detection
BullshitBench v2
95%Best
54%
Prompt injection robustness
Gray Swan IPI · k = 1
0.5%
13.4%Best
Prompt injection robustness
Gray Swan IPI · k = 10
4.1%
54.2%Best
Prompt injection robustness
Gray Swan IPI · k = 15
5.5%
60.8%Best
Agentic coding
SWE-Bench Pro
69.2%Best
64.7%
Multilingual coding
SWE-Bench Multilingual
84.4%Best
78%
Agentic coding
CursorBench v3.2
62.3%
66.7%Best
Agentic coding
CursorBench v3.1
63.8%
Agentic coding
DeepSWE 1.1
54%
Agentic coding
DeepSWE 1.0
55.8%
62%Best
Next.js coding
Next.js Evals
88%Best
83%
Agentic computer work
Frontier-Bench v0.1
21.1%Best
17.8%
Agentic terminal coding
Terminal-Bench 2.1
74.6%
83.3%Best
Browser agent
BU Bench
74%
Web browsing
BrowseComp
84.3%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
49.8%
Multidisciplinary reasoning
Humanity's Last Exam · with tools
57.9%
Agentic computer use
OSWorld-Verified
83.4%
Agentic financial analysis
Finance Agent v2
53.9%
Agentic legal work
Harvey's Legal Agent Benchmark
9.58%
12.92%Best
Tax questions
TaxEval v2
75.63%
Medical admin work
MedScribe
85.75%
86.88%Best
Knowledge work
GDPval-AA
1890
Knowledge work
GDPval-AA v2
1600
Community preference
Arena Elo (Text)
1482Best
1468
Community preference (code)
Arena Elo (Code)
1568Best
1549
Overview
CompanyAnthropicSpaceXAI
Release dateMay 28 2026Jul 8 2026
AccessProprietaryProprietary

Which is better: Claude Opus 4.8 or Grok 4.5?

Claude Opus 4.8 leads Grok 4.5 on 10 of the 15 benchmarks they both report. Claude Opus 4.8 shipped 41 days before Grok 4.5, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude Opus 4.8 leads at 95% vs Grok 4.5 at 54%. On Gray Swan IPI · k = 1, Claude Opus 4.8 leads at 0.5% vs Grok 4.5 at 13.4%. On Gray Swan IPI · k = 10, Claude Opus 4.8 leads at 4.1% vs Grok 4.5 at 54.2%. On Gray Swan IPI · k = 15, Claude Opus 4.8 leads at 5.5% vs Grok 4.5 at 60.8%. On SWE-Bench Pro, Claude Opus 4.8 leads at 69.2% vs Grok 4.5 at 64.7%. On SWE-Bench Multilingual, Claude Opus 4.8 leads at 84.4% vs Grok 4.5 at 78%. On CursorBench v3.2, Grok 4.5 leads at 66.7% vs Claude Opus 4.8 at 62.3%. On DeepSWE 1.0, Grok 4.5 leads at 62% vs Claude Opus 4.8 at 55.8%. On Next.js Evals, Claude Opus 4.8 leads at 88% vs Grok 4.5 at 83%. On Frontier-Bench v0.1, Claude Opus 4.8 leads at 21.1% vs Grok 4.5 at 17.8%. On Terminal-Bench 2.1, Grok 4.5 leads at 83.3% vs Claude Opus 4.8 at 74.6%. On Harvey's Legal Agent Benchmark, Grok 4.5 leads at 12.92% vs Claude Opus 4.8 at 9.58%. On MedScribe, Grok 4.5 leads at 86.88% vs Claude Opus 4.8 at 85.75%. On Arena Elo (Text), Claude Opus 4.8 leads at 1482 vs Grok 4.5 at 1468. On Arena Elo (Code), Claude Opus 4.8 leads at 1568 vs Grok 4.5 at 1549.

Frequently asked questions

Claude Opus 4.8 was released by Anthropic on May 28 2026.

Grok 4.5 was released by SpaceXAI on Jul 8 2026.

Claude Opus 4.8 leads on SWE-Bench Pro — Claude Opus 4.8 69.2% vs Grok 4.5 64.7%.

Other comparisons