Devstral SmallvsComposer 2.5
Devstral Small | Composer 2.5 | |
|---|---|---|
| Specifications | ||
ParametersA rough measure of how big the model is. More parameters usually means more capable and more expensive to run, though it is a poor guide on its own — a smaller, newer model often beats a larger, older one. | 24B | — |
| BenchmarksPublished by one model only | ||
SWE-Bench ProAgentic coding — Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better. | — | 54% |
SWE-Bench MultilingualMultilingual coding — Like SWE-Bench, but the coding problems span many programming languages, not just one. Tests how broadly the AI can code. Higher is better. | — | 79.8% |
CursorBench v3.2Agentic coding — Cursor's own test of harder, real-world coding tasks inside a code editor, on the refreshed v3.2 task set. Scores aren't comparable with v3.1. Higher is better. | — | 56.1% |
CursorBench v3.1Agentic coding — Cursor's own test of harder, real-world coding tasks inside a code editor. Higher is better. | — | 63.2% |
DeepSWE 1.0Agentic coding — Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. Higher is better. | — | 18% |
Next.js EvalsNext.js coding — Vercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better. | — | 92% |
Terminal-Bench 2.1Agentic terminal coding — Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better. | — | 73% |
Terminal-Bench 2.0Agentic terminal coding — Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? (Version 2.0 of the test.) Higher is better. | — | 69.3% |
| Overview | ||
| Company | Mistral | SpaceXAI |
| Release date | May 21 2025 | May 18 2026 |
| Access | Open Weight | Proprietary |
Other comparisons
Devstral SmallvsClaude Opus 5Composer 2.5vsClaude Opus 5Devstral SmallvsGPT-5.6 SolComposer 2.5vsGPT-5.6 SolDevstral SmallvsGemini 3.7 FlashComposer 2.5vsGemini 3.7 FlashDevstral SmallvsMuse GlimmerComposer 2.5vsMuse GlimmerDevstral SmallvsDeepSeek-V4-Pro-0813Composer 2.5vsDeepSeek-V4-Pro-0813Devstral SmallvsKimi K3Composer 2.5vsKimi K3
Frequently asked questions
Devstral Small and Composer 2.5 don't publish scores on any of the same benchmarks, so there's no direct head-to-head comparison. Devstral Small shipped 362 days before Composer 2.5, so benchmark comparisons should account for the intervening progress.
Devstral Small is open weight, while Composer 2.5 is proprietary.
Direct benchmark comparisons are unavailable — Devstral Small and Composer 2.5 don't publish scores on any of the same benchmarks.
Devstral Small was released by Mistral on May 21 2025.
Composer 2.5 was released by SpaceXAI on May 18 2026.
Devstral Small is an open weight model released by Mistral. Composer 2.5 is a proprietary model released by SpaceXAI.