GPT-5.4vsGrok 4.5

GPT-5.4
Grok 4.5
Benchmarks
Nonsense detection
BullshitBench v2
48%
54%Best
Prompt injection robustness
Gray Swan IPI · k = 1
13.4%
Prompt injection robustness
Gray Swan IPI · k = 10
54.2%
Prompt injection robustness
Gray Swan IPI · k = 15
60.8%
Agentic coding
SWE-Bench Pro
64.7%
Multilingual coding
SWE-Bench Multilingual
78%
Agentic coding
CursorBench v3.2
66.7%
Agentic coding
DeepSWE 1.1
54%
Agentic coding
DeepSWE 1.0
62%
Agentic coding
FrontierCode v1.1 (Extended) · extended split
56.6%
Expert software engineering
APEX-SWE
53.6%
Next.js coding
Next.js Evals
83%
83%
Agentic computer work
Frontier-Bench v0.1
17.8%
Agentic terminal coding
Terminal-Bench 3.0
15.7%
Agentic terminal coding
Terminal-Bench 2.1
83.3%
Agentic terminal coding
Terminal-Bench 2.0
75.1%
Software engineering
Expert-SWE (Internal)
68.5%
Expert agentic work
APEX-Agents
47.1%
General tool use
Toolathlon
54.6%
Web browsing
BrowseComp
82.7%
Cybersecurity
CyberGym
79%
Advanced math
FrontierMath · Tier 1–3
47.6%
Advanced math
FrontierMath · Tier 4
27.1%
Science
GPQA Diamond
92.8%
Agentic computer use
OSWorld-Verified
75%
Agentic legal work
Harvey's Legal Agent Benchmark
12.92%
Medical admin work
MedScribe
86.88%
Overall intelligence
AA Intelligence Index
56
Knowledge work
GDPval-AA v2
1526
Knowledge work
AA-Briefcase
1313
Knowledge work
GDPval (win/tie rate)
83%
Community preference
Arena Elo (Text)
1476Best
1468
Community preference (code)
Arena Elo (Code)
1462
1555Best
Overview
CompanyOpenAISpaceXAI
Release dateMar 5 2026Jul 8 2026
AccessProprietaryProprietary

Which is better: GPT-5.4 or Grok 4.5?

Grok 4.5 leads GPT-5.4 on 2 of the 4 benchmarks they both report (BullshitBench v2, Next.js Evals, Arena Elo (Text), Arena Elo (Code)). GPT-5.4 shipped 125 days before Grok 4.5, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Grok 4.5 leads at 54% vs GPT-5.4 at 48%. On Next.js Evals, both models score 83%. On Arena Elo (Text), GPT-5.4 leads at 1476 vs Grok 4.5 at 1468. On Arena Elo (Code), Grok 4.5 leads at 1555 vs GPT-5.4 at 1462.

Frequently asked questions

GPT-5.4 was released by OpenAI on Mar 5 2026.

Grok 4.5 was released by SpaceXAI on Jul 8 2026.

Other comparisons