Abstract reasoning
ARC-AGI-2
Puzzle-style tests of abstract reasoning and pattern-finding — the kind of thing people find easy but AIs often struggle with. Higher is better.
Rankings
Higher is better#195%
GPT-6 Astravia BenchLM
292.5%
GPT-5.6 Solvia BenchLM
390.4%
Claude Opus 5via BenchLM
490%
Claude Fable 5.1
584.6%
GPT-5.5
683.9%
GPT-5.6 Terravia BenchLM
783.3%
GPT-5.4-Provia BenchLM
877.1%
Gemini 3.1 Pro
975.8%
Claude Opus 4.7
1073.95%
GPT-5.4via BenchLM
1172.1%
Gemini 3.5 Flash
1272.08%
Claude Opus 4.8via BenchLM
1359.54%
GPT-5.6 Lunavia BenchLM
1458.3%
Claude Sonnet 4.6
1553.3%
Grok 4.20 Betavia BenchLM
1652.9%
GPT-5.2via BenchLM
1752.64%
Grok 4.5via BenchLM
1842.5%
Muse Sparkvia BenchLM
1933.6%
Gemini 3.0 Flash
2031.1%
Gemini 3.0 Provia BenchLM
2113.6%
Claude Sonnet 4.5via BenchLM
Scores marked “via” above are quoted with attribution from BenchLM, retrieved 7 September 2026.
ARC-AGI-2 — frequently asked questions
Puzzle-style tests of abstract reasoning and pattern-finding — the kind of thing people find easy but AIs often struggle with. Higher is better.