Agentic legal work

Harvey's Legal Agent Benchmark

Harvey's test of whether an AI agent can complete real legal work — drafting and reviewing documents, working with spreadsheets and presentations, and navigating files the way a lawyer's assistant would. Higher is better.

Rankings

Higher is better

Harvey's Legal Agent Benchmark — frequently asked questions

What is Harvey's Legal Agent Benchmark?
Harvey's test of whether an AI agent can complete real legal work — drafting and reviewing documents, working with spreadsheets and presentations, and navigating files the way a lawyer's assistant would. Higher is better.
Which AI model scores highest on Harvey's Legal Agent Benchmark?
Gemini 3.7 Flash by Google holds the best Harvey's Legal Agent Benchmark result among tracked models, at 90.7% (released Aug 13 2026). Higher scores are better on this benchmark.
What are the top 5 models on Harvey's Legal Agent Benchmark?
1. Gemini 3.7 Flash (Google) — 90.7%; 2. Muse Spark 1.1 (Meta) — 20%; 3. Grok 4.6 (SpaceXAI) — 15.8%; 4. Grok 4.5 (SpaceXAI) — 12.92%; 5. Claude Fable 5 (Anthropic) — 11.25%.
How many models have a published Harvey's Legal Agent Benchmark score?
7 tracked models have a published Harvey's Legal Agent Benchmark score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.