Abstract reasoning

ARC-AGI-2

Puzzle-style tests of abstract reasoning and pattern-finding — the kind of thing people find easy but AIs often struggle with. Higher is better.

Rankings

Higher is better

ARC-AGI-2 — frequently asked questions

What is ARC-AGI-2?
Puzzle-style tests of abstract reasoning and pattern-finding — the kind of thing people find easy but AIs often struggle with. Higher is better.
Which AI model scores highest on ARC-AGI-2?
GPT-5.5 by OpenAI holds the best ARC-AGI-2 result among tracked models, at 84.6% (released Apr 23 2026). Higher scores are better on this benchmark.
What are the top 5 models on ARC-AGI-2?
1. GPT-5.5 (OpenAI) — 84.6%; 2. GPT-5.4-Pro (OpenAI) — 83.3%; 3. Gemini 3.1 Pro (Google) — 77.1%; 4. Claude Opus 4.7 (Anthropic) — 75.8%; 5. Gemini 3.5 Flash (Google) — 72.1%.
How many models have a published ARC-AGI-2 score?
11 tracked models have a published ARC-AGI-2 score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks