Multidisciplinary reasoning

Humanity's Last Exam (Verified)

The re-checked edition of Humanity's Last Exam: the same extremely hard expert questions, minus the ones found to be flawed or wrongly answered. Scores on it run lower than on the original exam, so read the two as separate tests rather than a before-and-after. Higher is better.

Rankings

Higher is better

Humanity's Last Exam (Verified) — frequently asked questions

What is Humanity's Last Exam (Verified)?
The re-checked edition of Humanity's Last Exam: the same extremely hard expert questions, minus the ones found to be flawed or wrongly answered. Scores on it run lower than on the original exam, so read the two as separate tests rather than a before-and-after. Higher is better.
Which AI model scores highest on Humanity's Last Exam (Verified)?
Gemini 3.7 Flash by Google holds the best Humanity's Last Exam (Verified) result among tracked models, at 53.6% (released Aug 13 2026). Higher scores are better on this benchmark.
What are the top 1 models on Humanity's Last Exam (Verified)?
1. Gemini 3.7 Flash (Google) — 53.6%.
How many models have a published Humanity's Last Exam (Verified) score?
1 tracked model has a published Humanity's Last Exam (Verified) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks