Claude 3.5 HaikuvsGrok 4.5

Claude 3.5 Haiku
Grok 4.5
Benchmarks
Nonsense detection
BullshitBench v2
50%
54%Best
Prompt injection robustness
Gray Swan IPI · k = 1
13.4%
Prompt injection robustness
Gray Swan IPI · k = 10
54.2%
Prompt injection robustness
Gray Swan IPI · k = 15
60.8%
Agentic coding
SWE-Bench Pro
64.7%
Coding
SWE-Bench Verified
40.6%
Multilingual coding
SWE-Bench Multilingual
78%
Agentic coding
CursorBench v3.2
66.7%
Agentic coding
DeepSWE 1.1
54%
Agentic coding
DeepSWE 1.0
62%
Agentic coding
FrontierCode v1.1 (Extended) · extended split
56.6%
Expert software engineering
APEX-SWE
53.6%
Next.js coding
Next.js Evals
83%
Agentic computer work
Frontier-Bench v0.1
17.8%
Agentic terminal coding
Terminal-Bench 3.0
15.7%
Agentic terminal coding
Terminal-Bench 2.1
83.3%
Expert agentic work
APEX-Agents
47.1%
Science
GPQA Diamond
41.6%
Agentic legal work
Harvey's Legal Agent Benchmark
12.92%
Medical admin work
MedScribe
86.88%
Overall intelligence
AA Intelligence Index
56
Knowledge work
GDPval-AA v2
1526
Knowledge work
AA-Briefcase
1313
Community preference
Arena Elo (Text)
1468
Community preference (code)
Arena Elo (Code)
1555
Overview
CompanyAnthropicSpaceXAI
Release dateOct 22 2024Jul 8 2026
AccessProprietaryProprietary

Which is better: Claude 3.5 Haiku or Grok 4.5?

Grok 4.5 leads Claude 3.5 Haiku on 1 of the 1 benchmark they both report (BullshitBench v2). Claude 3.5 Haiku shipped 624 days before Grok 4.5, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Grok 4.5 leads at 54% vs Claude 3.5 Haiku at 50%.

Frequently asked questions

Claude 3.5 Haiku was released by Anthropic on Oct 22 2024.

Grok 4.5 was released by SpaceXAI on Jul 8 2026.

Other comparisons