Agentic legal work

Harvey's Legal Agent Benchmark (Held-out)

Harvey's test of whether an AI agent can complete real legal work, scored on a held-out set of tasks the model makers never see — making the numbers harder to game. Higher is better.

Rankings

Higher is better

Harvey's Legal Agent Benchmark (Held-out) — frequently asked questions

What is Harvey's Legal Agent Benchmark (Held-out)?
Harvey's test of whether an AI agent can complete real legal work, scored on a held-out set of tasks the model makers never see — making the numbers harder to game. Higher is better.
Which AI model scores highest on Harvey's Legal Agent Benchmark (Held-out)?
Claude Opus 5 by Anthropic holds the best Harvey's Legal Agent Benchmark (Held-out) result among tracked models, at 11.7% (released Jul 24 2026). Higher scores are better on this benchmark.
What are the top 1 models on Harvey's Legal Agent Benchmark (Held-out)?
1. Claude Opus 5 (Anthropic) — 11.7%.
How many models have a published Harvey's Legal Agent Benchmark (Held-out) score?
1 tracked model has a published Harvey's Legal Agent Benchmark (Held-out) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks