Agentic legal work
Harvey's Legal Agent Benchmark
Harvey's test of whether an AI agent can complete real legal work — drafting and reviewing documents, working with spreadsheets and presentations, and navigating files the way a lawyer's assistant would. Higher is better.
Rankings
Higher is betterHarvey's Legal Agent Benchmark — frequently asked questions
- What is Harvey's Legal Agent Benchmark?
- Harvey's test of whether an AI agent can complete real legal work — drafting and reviewing documents, working with spreadsheets and presentations, and navigating files the way a lawyer's assistant would. Higher is better.
- Which AI model scores highest on Harvey's Legal Agent Benchmark?
- Gemini 3.7 Flash by Google holds the best Harvey's Legal Agent Benchmark result among tracked models, at 90.7% (released Aug 13 2026). Higher scores are better on this benchmark.
- What are the top 5 models on Harvey's Legal Agent Benchmark?
- 1. Gemini 3.7 Flash (Google) — 90.7%; 2. Muse Spark 1.1 (Meta) — 20%; 3. Grok 4.6 (SpaceXAI) — 15.8%; 4. Grok 4.5 (SpaceXAI) — 12.92%; 5. Claude Fable 5 (Anthropic) — 11.25%.
- How many models have a published Harvey's Legal Agent Benchmark score?
- 7 tracked models have a published Harvey's Legal Agent Benchmark score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.