Coding
SWE-Bench Verified
Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better.
Rankings
Higher is better#195.5%
Claude Fable 5
287.6%
Claude Opus 4.7
380.9%
Claude Opus 4.5
480.8%
Claude Opus 4.6
580.6%
Gemini 3.1 Pro
680.2%
Kimi K2.6
780%
GPT-5.2
879.6%
Claude Sonnet 4.6
978%
Gemini 3.0 Flash
1077.8%
GLM-5
1177.6%
Mistral Medium 3.5
1277.4%
Muse Spark
1377.2%
Claude Sonnet 4.5
1476.8%
Kimi K2.5
1576.4%
Qwen3.5OW
1676.3%
GPT-5.1
1776.2%
Gemini 3.0 Pro
1874.5%
Claude Opus 4.1
1973.8%
GLM-4.7
2073.4%
Qwen3.6OW
2173.3%
Claude Haiku 4.5
2272.7%
Claude Sonnet 4
2372.5%
Claude Opus 4
2471.3%
Kimi K2 Thinking
2570.6%
Qwen3-Coder-NextOW
2668%
GLM-4.6
2765.8%
Kimi K2
2765.8%
Kimi K2 (0905)
2964.2%
GLM-4.5
3062.3%
Claude 3.7 Sonnet
3159.8%
GLM-4.5-Air
3259.6%
Gemini 2.5 Pro
3349%
Claude 3.5 Sonnet (upgraded)
3440.6%
Claude 3.5 Haiku
3533.4%
Claude 3.5 Sonnet
3633%
Claude 3 Opus
SWE-Bench Verified — frequently asked questions
- What is SWE-Bench Verified?
- Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better.
- Which AI model scores highest on SWE-Bench Verified?
- Claude Fable 5 by Anthropic holds the best SWE-Bench Verified result among tracked models, at 95.5% (released Jun 9 2026). Higher scores are better on this benchmark.
- What are the top 5 models on SWE-Bench Verified?
- 1. Claude Fable 5 (Anthropic) — 95.5%; 2. Claude Opus 4.7 (Anthropic) — 87.6%; 3. Claude Opus 4.5 (Anthropic) — 80.9%; 4. Claude Opus 4.6 (Anthropic) — 80.8%; 5. Gemini 3.1 Pro (Google) — 80.6%.
- How many models have a published SWE-Bench Verified score?
- 36 tracked models have a published SWE-Bench Verified score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
- What is the best open model on SWE-Bench Verified?
- Qwen3.5 by Qwen is the highest-ranked model with downloadable weights on SWE-Bench Verified, scoring 76.4% at rank 15 overall.