Multidisciplinary reasoning
Humanity's Last Examwith tools
Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better.
Rankings
Higher is better#167.7%
Claude Opus 5.5
265.6%
Claude Fable 5.1
364.5%
Claude Fable 5via BenchLM
364.5%
Claude Mythos 5via BenchLM
563.9%
DeepSeek-V4.1-Flash
663.6%
Claude Opus 5
762.5%
GLM-5.3OW
862.1%
Muse Spark 1.1
960%
DeepSeek-V4-Pro-0813
1058.7%
GPT-5.4-Provia BenchLM
1157.9%
Claude Opus 4.8
1257.4%
Claude Sonnet 5
1357.2%
GPT-5.5-Provia BenchLM
1456%
Kimi K3
1555.3%
GLM-5.3-FlashOW
1654.7%
Claude Opus 4.7
1654.7%
GLM-5.2via BenchLM
1853%
Claude Opus 4.6via BenchLM
1952.3%
GLM-5.1via BenchLM
2052.2%
GPT-5.5
2152.1%
GPT-5.4via BenchLM
2251.4%
Gemini 3.1 Pro
2350.4%
GLM-5
2350.4%
Muse Spark
2549%
Claude Sonnet 4.6via BenchLM
2643.6%
Qwen3.8-Maxvia BenchLM
2742.8%
GLM-4.7
2841.5%
GPT-5.4 minivia BenchLM
2941.4%
Qwen3.7-Maxvia BenchLM
3040.2%
Gemini 3.5 Flashvia BenchLM
3137.7%
GPT-5.4 nanovia BenchLM
3235.9%
Qwen3.8-Flash-NextOWvia BenchLM
3335%
Grok 4.3 Betavia BenchLM
3434.8%
DeepSeek-V4-Flash-0731via BenchLM
3534.7%
Kimi K2.6via BenchLM
3534.7%
Qwen3.7-Plusvia BenchLM
3730.8%
Claude Opus 4.5via BenchLM
3730.8%
Qwen3.8-27BOWvia BenchLM
3930.1%
Kimi K2.5via BenchLM
4028.8%
Qwen3.6-Plusvia BenchLM
4126.7%
Nemotron 3 UltraOSvia BenchLM
4219%
gpt-oss-120bOW
4318.8%
Gemini 2.5 Provia BenchLM
4417.3%
gpt-oss-20bOW
Scores marked “via” above are quoted with attribution from BenchLM, retrieved 22 September 2026.
Humanity's Last Exam — frequently asked questions
Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better.