GPT-5.4vsComposer 2.5

GPT-5.4
Composer 2.5
Benchmarks
Nonsense detection
BullshitBench v2
48%
Agentic coding
SWE-Bench Pro
54%
Multilingual coding
SWE-Bench Multilingual
79.8%
Agentic coding
CursorBench v3.2
56.1%
Agentic coding
CursorBench v3.1
63.2%
Agentic coding
DeepSWE 1.0
18%
Next.js coding
Next.js Evals
83%
92%Best
Agentic terminal coding
Terminal-Bench 2.1
73%
Agentic terminal coding
Terminal-Bench 2.0
75.1%Best
69.3%
Software engineering
Expert-SWE (Internal)
68.5%
General tool use
Toolathlon
54.6%
Web browsing
BrowseComp
82.7%
Cybersecurity
CyberGym
79%
Advanced math
FrontierMath · Tier 1–3
47.6%
Advanced math
FrontierMath · Tier 4
27.1%
Science
GPQA Diamond
92.8%
Agentic computer use
OSWorld-Verified
75%
Knowledge work
GDPval (win/tie rate)
83%
Community preference
Arena Elo (Text)
1476
Community preference (code)
Arena Elo (Code)
1462
Overview
CompanyOpenAISpaceXAI
Release dateMar 5 2026May 18 2026
AccessProprietaryProprietary

Which is better: GPT-5.4 or Composer 2.5?

GPT-5.4 and Composer 2.5 are evenly matched across the 2 benchmarks they both report (Next.js Evals, Terminal-Bench 2.0). GPT-5.4 shipped 74 days before Composer 2.5, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On Next.js Evals, Composer 2.5 leads at 92% vs GPT-5.4 at 83%. On Terminal-Bench 2.0, GPT-5.4 leads at 75.1% vs Composer 2.5 at 69.3%.

Frequently asked questions

GPT-5.4 was released by OpenAI on Mar 5 2026.

Composer 2.5 was released by SpaceXAI on May 18 2026.

Composer 2.5 leads on Next.js Evals — GPT-5.4 83% vs Composer 2.5 92%.

Other comparisons