Agentic coding
SWE-Bench Pro
Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better.
Rankings
Higher is better#181.2%
Claude Fable 5.1
280.3%
Claude Fable 5
379.2%
Claude Opus 5
469.2%
Claude Opus 4.8
567.7%
Qwen3.8-Max
664.7%
Grok 4.5
764.3%
Claude Opus 4.7
863.2%
Claude Sonnet 5
962.5%
Qwen3.8-Flash-NextOW
1062.1%
GLM-5.2
1161.7%
Qwen3.8-27BOW
1261.5%
Muse Spark 1.1
1360.6%
Qwen3.7-Max
1458.6%
GPT-5.5
1555.1%
Gemini 3.5 Flash
1655%
Muse Spark
1754.2%
Gemini 3.1 Pro
1754.2%
Gemini 3.5 Flash-Lite
1954%
Composer 2.5
2051.2%
Muse GlimmerOW
2149.6%
Gemini 3.0 Flash
2249.5%
Qwen3.6OW
2344.3%
Qwen3-Coder-NextOW
SWE-Bench Pro — frequently asked questions
Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better.