Community preference (documents)
Arena Elo (Documents)
Real people hand two anonymous AIs the same PDF or document and vote for whichever answers better. The votes become a chess-style Elo rating on arena.ai. Unlike a fixed document benchmark, the files are whatever people actually brought along. Higher is better.
Rankings
Higher is better#11520
Claude Opus 5
21510
Claude Opus 4.6
31504
Claude Fable 5
41498
Claude Opus 4.7
51485
GPT-5.5
61483
Claude Sonnet 4.6
71479
GPT-5.6 Terra
71479
GPT-5.6 Sol
91475
Claude Opus 4.8
101472
Muse Spark 1.1
111470
GPT-5.4
111470
Claude Sonnet 5
131465
Gemini 3.5 Flash
141462
GPT-5.6 Luna
151454
Grok 4.5
161451
Kimi K2.6
171445
Gemini 3.1 Pro
181443
Muse Spark
191440
Qwen3.7-Plus
201433
Gemini 3.0 Pro
211429
Kimi K2.5
221421
Gemini 2.5 Pro
231413
Gemini 3.0 Flash
241405
GPT-5.2
251401
GPT-5.1
Arena Elo (Documents) — frequently asked questions
- What is Arena Elo (Documents)?
- Real people hand two anonymous AIs the same PDF or document and vote for whichever answers better. The votes become a chess-style Elo rating on arena.ai. Unlike a fixed document benchmark, the files are whatever people actually brought along. Higher is better.
- Which AI model scores highest on Arena Elo (Documents)?
- Claude Opus 5 by Anthropic holds the best Arena Elo (Documents) result among tracked models, at 1520 (released Jul 24 2026). Higher scores are better on this benchmark.
- What are the top 5 models on Arena Elo (Documents)?
- 1. Claude Opus 5 (Anthropic) — 1520; 2. Claude Opus 4.6 (Anthropic) — 1510; 3. Claude Fable 5 (Anthropic) — 1504; 4. Claude Opus 4.7 (Anthropic) — 1498; 5. GPT-5.5 (OpenAI) — 1485.
- How many models have a published Arena Elo (Documents) score?
- 25 tracked models have a published Arena Elo (Documents) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.