Agentic coding
SWE-Bench Pro
Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better.
Rankings
Higher is better180.3%
Claude Fable 5
Anthropic · Jun 9 2026
269.2%
Claude Opus 4.8
Anthropic · May 28 2026
364.7%
Grok 4.5
SpaceXAI · Jul 8 2026
464.3%
Claude Opus 4.7
Anthropic · Apr 16 2026
563.2%
Claude Sonnet 5
Anthropic · Jun 30 2026
662.1%
GLM-5.2
Z.ai · Jun 16 2026
761.5%
Muse Spark 1.1
Meta · Jul 9 2026
858.6%
GPT-5.5
OpenAI · Apr 23 2026
955.1%
Gemini 3.5 Flash
Google · May 19 2026
1055%
Muse Spark
Meta · Apr 8 2026
1154.2%
Gemini 3.1 Pro
Google · Feb 19 2026
1254%
Composer 2.5
Cursor · May 18 2026
1349.6%
Gemini 3.0 Flash
Google · Dec 17 2025