Agentic coding
SWE-Bench Pro
Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better.
Rankings
Higher is better#180.3%
Claude Fable 5
269.2%
Claude Opus 4.8
367.7%
Qwen3.8-Max
464.7%
Grok 4.5
564.3%
Claude Opus 4.7
663.2%
Claude Sonnet 5
762.1%
GLM-5.2
861.5%
Muse Spark 1.1
960.6%
Qwen3.7-Max
1058.6%
GPT-5.5
1155.1%
Gemini 3.5 Flash
1255%
Muse Spark
1354.2%
Gemini 3.1 Pro
1354.2%
Gemini 3.5 Flash-Lite
1554%
Composer 2.5
1649.6%
Gemini 3.0 Flash
1749.5%
Qwen3.6OW
1844.3%
Qwen3-Coder-NextOW
SWE-Bench Pro — frequently asked questions
- What is SWE-Bench Pro?
- Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better.
- Which AI model scores highest on SWE-Bench Pro?
- Claude Fable 5 by Anthropic holds the best SWE-Bench Pro result among tracked models, at 80.3% (released Jun 9 2026). Higher scores are better on this benchmark.
- What are the top 5 models on SWE-Bench Pro?
- 1. Claude Fable 5 (Anthropic) — 80.3%; 2. Claude Opus 4.8 (Anthropic) — 69.2%; 3. Qwen3.8-Max (Qwen) — 67.7%; 4. Grok 4.5 (SpaceXAI) — 64.7%; 5. Claude Opus 4.7 (Anthropic) — 64.3%.
- How many models have a published SWE-Bench Pro score?
- 18 tracked models have a published SWE-Bench Pro score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
- What is the best open model on SWE-Bench Pro?
- Qwen3.6 by Qwen is the highest-ranked model with downloadable weights on SWE-Bench Pro, scoring 49.5% at rank 17 overall.