Multimodal coding
SWE-Bench Multimodal
Real bug reports that arrive with pictures attached — a screenshot, a mockup, a page rendering wrongly — so the AI has to read the image as well as the code to work out what to fix. Higher is better.
Rankings
Higher is betterSWE-Bench Multimodal — frequently asked questions
- What is SWE-Bench Multimodal?
- Real bug reports that arrive with pictures attached — a screenshot, a mockup, a page rendering wrongly — so the AI has to read the image as well as the code to work out what to fix. Higher is better.
- Which AI model scores highest on SWE-Bench Multimodal?
- Claude Opus 5 by Anthropic holds the best SWE-Bench Multimodal result among tracked models, at 59.4% (released Jul 24 2026). Higher scores are better on this benchmark.
- What are the top 3 models on SWE-Bench Multimodal?
- 1. Claude Opus 5 (Anthropic) — 59.4%; 2. Claude Fable 5.1 (Anthropic) — 54.7%; 3. Claude Fable 5 (Anthropic) — 54.1%.
- How many models have a published SWE-Bench Multimodal score?
- 3 tracked models have a published SWE-Bench Multimodal score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.