Composer 2.5
Composer 2.5 is an AI model released by Cursor (now part of SpaceXAI) on Monday, May 18 2026, 31 days after Grok 4.3 Beta. Benchmark results (shown below) cover Next.js Evals, SWE-Bench Pro, SWE-Bench Multilingual, CursorBench v3.2, CursorBench v3.1, DeepSWE 1.0, and 2 more.
Benchmarks
Next.js coding
Next.js EvalsVercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better.
92%
#1Best published Next.js score of all tracked models
Agentic coding
SWE-Bench ProCan the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better.
54%
#16 of 20Best: Claude Fable 5 · 80.3%
Multilingual coding
SWE-Bench MultilingualLike SWE-Bench, but the coding problems span many programming languages, not just one. Tests how broadly the AI can code. Higher is better.
79.8%
#3 of 10Best: Claude Opus 4.8 · 84.4%
Agentic coding
CursorBench v3.2Cursor's own test of harder, real-world coding tasks inside a code editor, on the refreshed v3.2 task set. Scores aren't comparable with v3.1. Higher is better.
56.1%
#11 of 15Best: Claude Fable 5 · 70.5%
Agentic coding
CursorBench v3.1Cursor's own test of harder, real-world coding tasks inside a code editor. Higher is better.
63.2%
#5 of 12Best: Claude Fable 5 · 72.9%
Agentic coding
DeepSWE 1.0Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. Higher is better.
18%
#6 of 6Best: Kimi K3 · 67.5%
Agentic terminal coding
Terminal-Bench 2.1Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better.
73%
#18 of 25Best: GPT-5.6 Sol · 88.8%
Agentic terminal coding
Terminal-Bench 2.0Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? (Version 2.0 of the test.) Higher is better.
69.3%
#5 of 13Best: GPT-5.5 · 82.7%
Composer 2.5 — frequently asked questions
- When was Composer 2.5 released?
- Composer 2.5 was released by Cursor (now part of SpaceXAI) on Monday, May 18 2026.
- Who made Composer 2.5?
- Composer 2.5 was built by Cursor (now part of SpaceXAI). Elon Musk's AI company building the Grok series of models. Founded in 2023 as xAI, now part of SpaceX. Cursor (Anysphere) and its Composer coding models were acquired in 2026 and are tracked here.
- What benchmark scores did Composer 2.5 get?
- Composer 2.5 reports 8 tracked benchmark scores — SWE-Bench Pro: 54%; SWE-Bench Multilingual: 79.8%; CursorBench v3.2: 56.1%; CursorBench v3.1: 63.2%; DeepSWE 1.0: 18%; Next.js Evals: 92%; Terminal-Bench 2.1: 73%; Terminal-Bench 2.0: 69.3%. Scores are the figures published at release by Cursor (now part of SpaceXAI). It holds the best score among all models tracked here on Next.js Evals.
- Is Composer 2.5 open source?
- No. Composer 2.5 is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.
- What came before and after Composer 2.5?
- SpaceXAI's previous tracked release was Grok 4.3 Beta on Apr 17 2026, 31 days earlier. It was followed by Grok 4.5 on Jul 8 2026.