Claude Sonnet 4.6vsGrok 4.5

Claude Sonnet 4.6
Grok 4.5
Benchmarks
Nonsense detection
BullshitBench v2
91%Best
54%
Prompt injection robustness
Gray Swan IPI · k = 1
13.4%
Prompt injection robustness
Gray Swan IPI · k = 10
54.2%
Prompt injection robustness
Gray Swan IPI · k = 15
60.8%
Agentic coding
SWE-Bench Pro
64.7%
Coding
SWE-Bench Verified
79.6%
Multilingual coding
SWE-Bench Multilingual
78%
Agentic coding
CursorBench v3.2
66.7%
Agentic coding
CursorBench v3.1
49%
Agentic coding
DeepSWE 1.1
30%
54%Best
Agentic coding
DeepSWE 1.0
62%
Next.js coding
Next.js Evals
58%
83%Best
Agentic computer work
Frontier-Bench v0.1
17.8%
Agentic terminal coding
Terminal-Bench 2.1
83.3%
Multi-step tool use
MCP Atlas
69.5%
Browser agent
BU Bench
62%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
33.2%
Abstract reasoning
ARC-AGI-2
58.3%
Science
GPQA Diamond
89.9%
Agentic computer use
OSWorld-Verified
72.5%
Agentic financial analysis
Finance Agent v2
51%
Agentic legal work
Harvey's Legal Agent Benchmark
12.92%
Medical admin work
MedScribe
86.88%
Knowledge work
GDPval-AA
1676
Chart reasoning
CharXiv Reasoning
72.4%
Multimodal reasoning
MMMU-Pro
74.5%
Spatial reasoning
Blueprint-Bench 2
6.7%
Long context
MRCR v2 (8-needle) · 128k average
84.9%
Community preference (code)
Arena Elo (Code)
1521
1549Best
Overview
CompanyAnthropicSpaceXAI
Release dateFeb 17 2026Jul 8 2026
AccessProprietaryProprietary

Which is better: Claude Sonnet 4.6 or Grok 4.5?

Grok 4.5 leads Claude Sonnet 4.6 on 3 of the 4 benchmarks they both report (BullshitBench v2, DeepSWE 1.1, Next.js Evals, Arena Elo (Code)). Claude Sonnet 4.6 shipped 141 days before Grok 4.5, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude Sonnet 4.6 leads at 91% vs Grok 4.5 at 54%. On DeepSWE 1.1, Grok 4.5 leads at 54% vs Claude Sonnet 4.6 at 30%. On Next.js Evals, Grok 4.5 leads at 83% vs Claude Sonnet 4.6 at 58%. On Arena Elo (Code), Grok 4.5 leads at 1549 vs Claude Sonnet 4.6 at 1521.

Frequently asked questions

Claude Sonnet 4.6 was released by Anthropic on Feb 17 2026.

Grok 4.5 was released by SpaceXAI on Jul 8 2026.

Grok 4.5 leads on DeepSWE 1.1 — Claude Sonnet 4.6 30% vs Grok 4.5 54%.

Other comparisons