Multidisciplinary reasoning
Humanity's Last Examwith tools
Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better.
Rankings
Higher is betterHumanity's Last Exam — frequently asked questions
- What is Humanity's Last Exam?
- Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better.
- Which AI model scores highest on Humanity's Last Exam?
- Claude Opus 5 by Anthropic holds the best Humanity's Last Exam (with tools) result among tracked models, at 64.7% (released Jul 24 2026). Higher scores are better on this benchmark.
- What are the top 5 models on Humanity's Last Exam?
- 1. Claude Opus 5 (Anthropic) — 64.7%; 2. Claude Fable 5 (Anthropic) — 64.5%; 3. Muse Spark 1.1 (Meta) — 62.1%; 4. GPT-5.4-Pro (OpenAI) — 58.7%; 5. Claude Opus 4.8 (Anthropic) — 57.9%.
- How many models have a published Humanity's Last Exam score?
- 17 tracked models have a published Humanity's Last Exam (with tools) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.