# Humanity's Last Exam (no tools) — AI model rankings

Humanity's Last Exam — extremely hard expert questions across many subjects, written so you can't just look up the answer. “No tools” means the AI answers on its own. Higher is better.

22 tracked models have a published Humanity's Last Exam (no tools) score. Higher is better. Scores come from published lab reports and benchmark sources; recorded sources appear beside the scores.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | Claude Fable 5.1 | Anthropic | 60.9% | Lab | Sep 1 2026 |
| 2 | Claude Opus 5 | Anthropic | 56.3% | Lab | Jul 24 2026 |
| 3 | Claude Opus 4.8 | Anthropic | 49.8% | Lab | May 28 2026 |
| 4 | Claude Opus 4.7 | Anthropic | 46.9% | Lab | Apr 16 2026 |
| 5 | Gemini 3.1 Pro | Google | 44.4% | Lab | Feb 19 2026 |
| 6 | Kimi K3 | Moonshot AI | 43.5% | Lab | Jul 16 2026 |
| 7 | Claude Sonnet 5 | Anthropic | 43.2% | Lab | Jun 30 2026 |
| 8 | DeepSeek-V4-Pro-0813 | DeepSeek | 42.7% | Lab | Aug 13 2026 |
| 9 | GPT-5.5 | OpenAI | 41.4% | Lab | Apr 23 2026 |
| 10 | GLM-5.2 | Z.ai | 40.5% | Lab | Jun 16 2026 |
| 11 | Gemini 3.5 Flash | Google | 40.2% | Lab | May 19 2026 |
| 12 | DeepSeek-V4.1-Flash | DeepSeek | 36.8% | Lab | Sep 10 2026 |
| 13 | Qwen3.8-Flash-Next | Qwen | 35.9% | Lab | Aug 26 2026 |
| 14 | Gemini 3.0 Flash | Google | 33.7% | Lab | Dec 17 2025 |
| 15 | Claude Sonnet 4.6 | Anthropic | 33.2% | Lab | Feb 17 2026 |
| 16 | Qwen3.8-27B | Qwen | 30.8% | Lab | Aug 14 2026 |
| 17 | Qwen3.5 | Qwen | 28.7% | Lab | Feb 16 2026 |
| 18 | GLM-4.7 | Z.ai | 24.8% | Lab | Dec 22 2025 |
| 19 | Muse Glimmer | Meta | 22% | Lab | Aug 10 2026 |
| 20 | Qwen3.6 | Qwen | 21.4% | Lab | Apr 16 2026 |
| 21 | gpt-oss-120b | OpenAI | 14.9% | Lab | Aug 5 2025 |
| 22 | gpt-oss-20b | OpenAI | 10.9% | Lab | Aug 5 2025 |


---

Canonical page: https://aireleasetracker.com/benchmark/hle-no-tools
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
