Multilingual coding
SWE-Bench Multilingual
Like SWE-Bench, but the coding problems span many programming languages, not just one. Tests how broadly the AI can code. Higher is better.
Rankings
Higher is betterSWE-Bench Multilingual — frequently asked questions
- What is SWE-Bench Multilingual?
- Like SWE-Bench, but the coding problems span many programming languages, not just one. Tests how broadly the AI can code. Higher is better.
- Which AI model scores highest on SWE-Bench Multilingual?
- Claude Opus 4.8 by Anthropic holds the best SWE-Bench Multilingual result among tracked models, at 84.4% (released May 28 2026). Higher scores are better on this benchmark.
- What are the top 5 models on SWE-Bench Multilingual?
- 1. Claude Opus 4.8 (Anthropic) — 84.4%; 2. Claude Opus 4.7 (Anthropic) — 80.5%; 3. Composer 2.5 (Cursor) — 79.8%; 4. Grok 4.5 (SpaceXAI) — 78%; 5. GPT-5.5 (OpenAI) — 77.8%.
- How many models have a published SWE-Bench Multilingual score?
- 10 tracked models have a published SWE-Bench Multilingual score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
- What is the best open model on SWE-Bench Multilingual?
- Qwen3.5 by Qwen is the highest-ranked model with downloadable weights on SWE-Bench Multilingual, scoring 69.3% at rank 8 overall.