Knowledge work
GDPval-AA
Measures how well the AI does economically valuable knowledge work, judged against human experts. Shown as a rating (like a chess Elo) — higher is better.
Rankings
Higher is betterGDPval-AA — frequently asked questions
- What is GDPval-AA?
- Measures how well the AI does economically valuable knowledge work, judged against human experts. Shown as a rating (like a chess Elo) — higher is better.
- Which AI model scores highest on GDPval-AA?
- Claude Fable 5 by Anthropic holds the best GDPval-AA result among tracked models, at 1932 (released Jun 9 2026). Higher scores are better on this benchmark.
- What are the top 5 models on GDPval-AA?
- 1. Claude Fable 5 (Anthropic) — 1932; 2. Claude Opus 4.8 (Anthropic) — 1890; 3. GPT-5.5 (OpenAI) — 1769; 4. Claude Opus 4.7 (Anthropic) — 1753; 5. Claude Sonnet 4.6 (Anthropic) — 1676.
- How many models have a published GDPval-AA score?
- 9 tracked models have a published GDPval-AA score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.