# Humanity's Last Exam (with tools) — AI model rankings

Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better.

17 tracked models have a published Humanity's Last Exam (with tools) score. Higher is better. Scores are as published at each model's release.

## Ranking

| Rank | Model | Developer | Score | Released |
| --- | --- | --- | --- | --- |
| 1 | Claude Opus 5 | Anthropic | 64.7% | Jul 24 2026 |
| 2 | Claude Fable 5 | Anthropic | 64.5% | Jun 9 2026 |
| 3 | Muse Spark 1.1 | Meta | 62.1% | Jul 9 2026 |
| 4 | GPT-5.4-Pro | OpenAI | 58.7% | Mar 5 2026 |
| 5 | Claude Opus 4.8 | Anthropic | 57.9% | May 28 2026 |
| 6 | Claude Sonnet 5 | Anthropic | 57.4% | Jun 30 2026 |
| 7 | GPT-5.5-Pro | OpenAI | 57.2% | Apr 23 2026 |
| 8 | Kimi K3 | Moonshot AI | 56% | Jul 16 2026 |
| 9 | Claude Opus 4.7 | Anthropic | 54.7% | Apr 16 2026 |
| 9 | GLM-5.2 | Z.ai | 54.7% | Jun 16 2026 |
| 11 | Claude Opus 4.6 | Anthropic | 53% | Feb 5 2026 |
| 12 | GLM-5.1 | Z.ai | 52.3% | Apr 7 2026 |
| 13 | GPT-5.5 | OpenAI | 52.2% | Apr 23 2026 |
| 14 | Gemini 3.1 Pro | Google | 51.4% | Feb 19 2026 |
| 15 | GLM-5 | Z.ai | 50.4% | Feb 12 2026 |
| 15 | Muse Spark | Meta | 50.4% | Apr 8 2026 |
| 17 | GLM-4.7 | Z.ai | 42.8% | Dec 22 2025 |


---

Canonical page: https://aireleasetracker.com/benchmark/hle-with-tools
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
