# SWE-Bench Multimodal — AI model rankings

Real bug reports that arrive with pictures attached — a screenshot, a mockup, a page rendering wrongly — so the AI has to read the image as well as the code to work out what to fix. Higher is better.

3 tracked models have a published SWE-Bench Multimodal score. Higher is better. Scores are as published at each model's release.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 5 | Anthropic | 59.4% | Lab | Jul 24 2026 |
| 2 | Claude Fable 5.1 | Anthropic | 54.7% | Lab | Sep 1 2026 |
| 3 | Claude Fable 5 | Anthropic | 54.1% | Lab | Jun 9 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/swe-bench-multimodal
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
