Multidisciplinary reasoning

Humanity's Last Examno tools

Humanity's Last Exam — extremely hard expert questions across many subjects, written so you can't just look up the answer. “No tools” means the AI answers on its own. Higher is better.

Rankings

Higher is better

Humanity's Last Exam — frequently asked questions

What is Humanity's Last Exam?
Humanity's Last Exam — extremely hard expert questions across many subjects, written so you can't just look up the answer. “No tools” means the AI answers on its own. Higher is better.
Which AI model scores highest on Humanity's Last Exam?
Claude Opus 5 by Anthropic holds the best Humanity's Last Exam (no tools) result among tracked models, at 56.3% (released Jul 24 2026). Higher scores are better on this benchmark.
What are the top 5 models on Humanity's Last Exam?
1. Claude Opus 5 (Anthropic) — 56.3%; 2. Claude Opus 4.8 (Anthropic) — 49.8%; 3. Claude Opus 4.7 (Anthropic) — 46.9%; 4. Gemini 3.1 Pro (Google) — 44.4%; 5. Kimi K3 (Moonshot AI) — 43.5%.
How many models have a published Humanity's Last Exam score?
14 tracked models have a published Humanity's Last Exam (no tools) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
What is the best open model on Humanity's Last Exam?
Qwen3.5 by Qwen is the highest-ranked model with downloadable weights on Humanity's Last Exam (no tools), scoring 28.7% at rank 12 overall.
← All benchmarks