Expert software engineering

APEX-SWE

expert-level software-engineering tasks (AI Productivity Index)

Rankings

Higher is better

APEX-SWE — frequently asked questions

What is APEX-SWE?
expert-level software-engineering tasks (AI Productivity Index)
Which AI model scores highest on APEX-SWE?
Grok 4.6 by SpaceXAI holds the best APEX-SWE result among tracked models, at 56.4% (released Aug 12 2026). Higher scores are better on this benchmark.
What are the top 2 models on APEX-SWE?
1. Grok 4.6 (SpaceXAI) — 56.4%; 2. Grok 4.5 (SpaceXAI) — 53.6%.
How many models have a published APEX-SWE score?
2 tracked models have a published APEX-SWE score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks