Multidisciplinary reasoning
Humanity's Last Examno tools
Humanity's Last Exam — extremely hard expert questions across many subjects, written so you can't just look up the answer. “No tools” means the AI answers on its own. Higher is better.
Rankings
Higher is better#160.9%
Claude Fable 5.1
256.3%
Claude Opus 5
349.8%
Claude Opus 4.8
446.9%
Claude Opus 4.7
544.4%
Gemini 3.1 Pro
643.5%
Kimi K3
743.2%
Claude Sonnet 5
842.7%
DeepSeek-V4-Pro-0813OW
941.4%
GPT-5.5
1040.5%
GLM-5.2
1140.2%
Gemini 3.5 Flash
1236.8%
DeepSeek-V4.1-FlashOW
1335.9%
Qwen3.8-Flash-NextOW
1433.7%
Gemini 3.0 Flash
1533.2%
Claude Sonnet 4.6
1630.8%
Qwen3.8-27BOW
1728.7%
Qwen3.5OW
1824.8%
GLM-4.7
1922%
Muse GlimmerOW
2021.4%
Qwen3.6OW
2114.9%
gpt-oss-120bOW
2210.9%
gpt-oss-20bOW
Humanity's Last Exam — frequently asked questions
Humanity's Last Exam — extremely hard expert questions across many subjects, written so you can't just look up the answer. “No tools” means the AI answers on its own. Higher is better.