Agentic legal work
Harvey's Legal Agent Benchmark (Held-out)
Harvey's test of whether an AI agent can complete real legal work, scored on a held-out set of tasks the model makers never see — making the numbers harder to game. Higher is better.
Rankings
Higher is betterHarvey's Legal Agent Benchmark (Held-out) — frequently asked questions
- What is Harvey's Legal Agent Benchmark (Held-out)?
- Harvey's test of whether an AI agent can complete real legal work, scored on a held-out set of tasks the model makers never see — making the numbers harder to game. Higher is better.
- Which AI model scores highest on Harvey's Legal Agent Benchmark (Held-out)?
- Claude Opus 5 by Anthropic holds the best Harvey's Legal Agent Benchmark (Held-out) result among tracked models, at 11.7% (released Jul 24 2026). Higher scores are better on this benchmark.
- What are the top 1 models on Harvey's Legal Agent Benchmark (Held-out)?
- 1. Claude Opus 5 (Anthropic) — 11.7%.
- How many models have a published Harvey's Legal Agent Benchmark (Held-out) score?
- 1 tracked model has a published Harvey's Legal Agent Benchmark (Held-out) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.