Agentic coding

SWE-Bench Pro

Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better.

Rankings

Higher is better

SWE-Bench Pro — frequently asked questions

What is SWE-Bench Pro?
Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better.
Which AI model scores highest on SWE-Bench Pro?
Claude Fable 5 by Anthropic holds the best SWE-Bench Pro result among tracked models, at 80.3% (released Jun 9 2026). Higher scores are better on this benchmark.
What are the top 5 models on SWE-Bench Pro?
1. Claude Fable 5 (Anthropic) — 80.3%; 2. Claude Opus 4.8 (Anthropic) — 69.2%; 3. Qwen3.8-Max (Qwen) — 67.7%; 4. Grok 4.5 (SpaceXAI) — 64.7%; 5. Claude Opus 4.7 (Anthropic) — 64.3%.
How many models have a published SWE-Bench Pro score?
18 tracked models have a published SWE-Bench Pro score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
What is the best open model on SWE-Bench Pro?
Qwen3.6 by Qwen is the highest-ranked model with downloadable weights on SWE-Bench Pro, scoring 49.5% at rank 17 overall.
← All benchmarks