Multilingual coding

SWE-Bench Multilingual

Like SWE-Bench, but the coding problems span many programming languages, not just one. Tests how broadly the AI can code. Higher is better.

Rankings

Higher is better

SWE-Bench Multilingual — frequently asked questions

What is SWE-Bench Multilingual?
Like SWE-Bench, but the coding problems span many programming languages, not just one. Tests how broadly the AI can code. Higher is better.
Which AI model scores highest on SWE-Bench Multilingual?
Claude Opus 4.8 by Anthropic holds the best SWE-Bench Multilingual result among tracked models, at 84.4% (released May 28 2026). Higher scores are better on this benchmark.
What are the top 5 models on SWE-Bench Multilingual?
1. Claude Opus 4.8 (Anthropic) — 84.4%; 2. Claude Opus 4.7 (Anthropic) — 80.5%; 3. Composer 2.5 (Cursor) — 79.8%; 4. Grok 4.5 (SpaceXAI) — 78%; 5. GPT-5.5 (OpenAI) — 77.8%.
How many models have a published SWE-Bench Multilingual score?
10 tracked models have a published SWE-Bench Multilingual score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
What is the best open model on SWE-Bench Multilingual?
Qwen3.5 by Qwen is the highest-ranked model with downloadable weights on SWE-Bench Multilingual, scoring 69.3% at rank 8 overall.
← All benchmarks