Knowledge work

GDPval-AA

Measures how well the AI does economically valuable knowledge work, judged against human experts. Shown as a rating (like a chess Elo) — higher is better.

Rankings

Higher is better

GDPval-AA — frequently asked questions

What is GDPval-AA?
Measures how well the AI does economically valuable knowledge work, judged against human experts. Shown as a rating (like a chess Elo) — higher is better.
Which AI model scores highest on GDPval-AA?
Claude Fable 5 by Anthropic holds the best GDPval-AA result among tracked models, at 1932 (released Jun 9 2026). Higher scores are better on this benchmark.
What are the top 5 models on GDPval-AA?
1. Claude Fable 5 (Anthropic) — 1932; 2. Claude Opus 4.8 (Anthropic) — 1890; 3. GPT-5.5 (OpenAI) — 1769; 4. Claude Opus 4.7 (Anthropic) — 1753; 5. Claude Sonnet 4.6 (Anthropic) — 1676.
How many models have a published GDPval-AA score?
9 tracked models have a published GDPval-AA score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks