Advanced math

FrontierMathTier 4

Very hard, research-level math problems. Tier 4 is the hardest — close to what professional research mathematicians tackle. Higher is better.

Rankings

Higher is better

FrontierMath — frequently asked questions

What is FrontierMath?
Very hard, research-level math problems. Tier 4 is the hardest — close to what professional research mathematicians tackle. Higher is better.
Which AI model scores highest on FrontierMath?
GPT-5.5-Pro by OpenAI holds the best FrontierMath (Tier 4) result among tracked models, at 39.6% (released Apr 23 2026). Higher scores are better on this benchmark.
What are the top 5 models on FrontierMath?
1. GPT-5.5-Pro (OpenAI) — 39.6%; 2. GPT-5.4-Pro (OpenAI) — 38%; 3. GPT-5.5 (OpenAI) — 35.4%; 4. GPT-5.4 (OpenAI) — 27.1%; 5. Claude Opus 4.7 (Anthropic) — 22.9%.
How many models have a published FrontierMath score?
6 tracked models have a published FrontierMath (Tier 4) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks