Composer 2.5vsGrok 4.6

Composer 2.5
Grok 4.6
Benchmarks
Agentic coding
SWE-Bench Pro
54%
Multilingual coding
SWE-Bench Multilingual
79.8%
Agentic coding
CursorBench v3.2
56.1%
69.9%Best
Agentic coding
CursorBench v3.1
63.2%
Agentic coding
DeepSWE 1.1
65.9%
Agentic coding
DeepSWE 1.0
18%
Agentic coding
FrontierCode v1.1 (Extended) · extended split
61.3%
Expert software engineering
APEX-SWE
56.4%
Next.js coding
Next.js Evals
92%
Agentic terminal coding
Terminal-Bench 3.0
26%
Agentic terminal coding
Terminal-Bench 2.1
73%
Agentic terminal coding
Terminal-Bench 2.0
69.3%
Expert agentic work
APEX-Agents
57.5%
Agentic legal work
Harvey's Legal Agent Benchmark
15.8%
Overall intelligence
AA Intelligence Index
61
Knowledge work
GDPval-AA v2
1753
Knowledge work
AA-Briefcase
1577
Overview
CompanyCursorSpaceXAI
Release dateMay 18 2026Aug 12 2026
AccessProprietaryProprietary

Which is better: Composer 2.5 or Grok 4.6?

Grok 4.6 leads Composer 2.5 on 1 of the 1 benchmark they both report (CursorBench v3.2). Composer 2.5 shipped 86 days before Grok 4.6, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On CursorBench v3.2, Grok 4.6 leads at 69.9% vs Composer 2.5 at 56.1%.

Frequently asked questions

Composer 2.5 was released by Cursor on May 18 2026.

Grok 4.6 was released by SpaceXAI on Aug 12 2026.

Grok 4.6 leads on CursorBench v3.2 — Composer 2.5 56.1% vs Grok 4.6 69.9%.

Other comparisons