GPT-5.6 LunavsGrok 4.5

GPT-5.6 Luna
Grok 4.5
Benchmarks
Nonsense detection
BullshitBench v2
40%
54%Best
Prompt injection robustness
Gray Swan IPI · k = 1
8.3%
13.4%Best
Prompt injection robustness
Gray Swan IPI · k = 10
38.6%
54.2%Best
Prompt injection robustness
Gray Swan IPI · k = 15
43.9%
60.8%Best
Agentic coding
SWE-Bench Pro
64.7%
Multilingual coding
SWE-Bench Multilingual
78%
Agentic coding
CursorBench v3.2
61.1%
66.7%Best
Agentic coding
DeepSWE 1.1
54%
Agentic coding
DeepSWE 1.0
62%
Next.js coding
Next.js Evals
83%
Agentic computer work
Frontier-Bench v0.1
14.3%
17.8%Best
Agentic terminal coding
Terminal-Bench 2.1
82.5%
83.3%Best
Agentic legal work
Harvey's Legal Agent Benchmark
12.92%
Medical admin work
MedScribe
86.88%
Community preference
Arena Elo (Text)
1468
Community preference (code)
Arena Elo (Code)
1523
1549Best
Overview
CompanyOpenAISpaceXAI
Release dateJun 26 2026Jul 8 2026
AccessProprietaryProprietary

Which is better: GPT-5.6 Luna or Grok 4.5?

Grok 4.5 leads GPT-5.6 Luna on 5 of the 8 benchmarks they both report. GPT-5.6 Luna shipped 12 days before Grok 4.5, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Grok 4.5 leads at 54% vs GPT-5.6 Luna at 40%. On Gray Swan IPI · k = 1, GPT-5.6 Luna leads at 8.3% vs Grok 4.5 at 13.4%. On Gray Swan IPI · k = 10, GPT-5.6 Luna leads at 38.6% vs Grok 4.5 at 54.2%. On Gray Swan IPI · k = 15, GPT-5.6 Luna leads at 43.9% vs Grok 4.5 at 60.8%. On CursorBench v3.2, Grok 4.5 leads at 66.7% vs GPT-5.6 Luna at 61.1%. On Frontier-Bench v0.1, Grok 4.5 leads at 17.8% vs GPT-5.6 Luna at 14.3%. On Terminal-Bench 2.1, Grok 4.5 leads at 83.3% vs GPT-5.6 Luna at 82.5%. On Arena Elo (Code), Grok 4.5 leads at 1549 vs GPT-5.6 Luna at 1523.

Frequently asked questions

GPT-5.6 Luna was released by OpenAI on Jun 26 2026.

Grok 4.5 was released by SpaceXAI on Jul 8 2026.

Grok 4.5 leads on CursorBench v3.2 — GPT-5.6 Luna 61.1% vs Grok 4.5 66.7%.

Other comparisons