Coding

SWE-Bench Verified

Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better.

Rankings

Higher is better

SWE-Bench Verified — frequently asked questions

What is SWE-Bench Verified?
Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better.
Which AI model scores highest on SWE-Bench Verified?
Claude Fable 5 by Anthropic holds the best SWE-Bench Verified result among tracked models, at 95.5% (released Jun 9 2026). Higher scores are better on this benchmark.
What are the top 5 models on SWE-Bench Verified?
1. Claude Fable 5 (Anthropic) — 95.5%; 2. Claude Opus 4.7 (Anthropic) — 87.6%; 3. Claude Opus 4.5 (Anthropic) — 80.9%; 4. Claude Opus 4.6 (Anthropic) — 80.8%; 5. Gemini 3.1 Pro (Google) — 80.6%.
How many models have a published SWE-Bench Verified score?
36 tracked models have a published SWE-Bench Verified score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
What is the best open model on SWE-Bench Verified?
Qwen3.5 by Qwen is the highest-ranked model with downloadable weights on SWE-Bench Verified, scoring 76.4% at rank 15 overall.
← All benchmarks