Claude Sonnet 5vsGrok 4.1

Claude Sonnet 5
Grok 4.1
Benchmarks
Nonsense detection
BullshitBench v2
80%
Prompt injection robustness
Gray Swan IPI · k = 1
0.6%
Prompt injection robustness
Gray Swan IPI · k = 10
4.7%
Prompt injection robustness
Gray Swan IPI · k = 15
5.9%
Agentic coding
SWE-Bench Pro
63.2%
Agentic coding
CursorBench v3.2
61.5%
Agentic coding
CursorBench v3.1
61.2%
Next.js coding
Next.js Evals
79%
Supabase coding
Supabase Evals · with skills
95.5%
Supabase coding
Supabase Evals · no skills
90.9%
Agentic computer work
Frontier-Bench v0.1
14.6%
Agentic terminal coding
Terminal-Bench 2.1
80.4%
Web browsing
BrowseComp
84.7%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
43.2%
Multidisciplinary reasoning
Humanity's Last Exam · with tools
57.4%
Agentic computer use
OSWorld-Verified
81.2%
Knowledge work
GDPval-AA
1618
Community preference
Arena Elo (Text)
1463
1466Best
Community preference (code)
Arena Elo (Code)
1543Best
1209
Overview
CompanyAnthropicSpaceXAI
Release dateJun 30 2026Nov 17 2025
AccessProprietaryProprietary

Which is better: Claude Sonnet 5 or Grok 4.1?

Claude Sonnet 5 and Grok 4.1 are evenly matched across the 2 benchmarks they both report (Arena Elo (Text), Arena Elo (Code)). Grok 4.1 shipped 225 days before Claude Sonnet 5, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On Arena Elo (Text), Grok 4.1 leads at 1466 vs Claude Sonnet 5 at 1463. On Arena Elo (Code), Claude Sonnet 5 leads at 1543 vs Grok 4.1 at 1209.

Frequently asked questions

Claude Sonnet 5 was released by Anthropic on Jun 30 2026.

Grok 4.1 was released by SpaceXAI on Nov 17 2025.

Other comparisons