Long context
MRCR v2 (8-needle)128k average
Tests whether the AI can find specific details buried inside a very long document (around 128k tokens — roughly a long book). Higher is better.
Rankings
Higher is betterMRCR v2 (8-needle) — frequently asked questions
- What is MRCR v2 (8-needle)?
- Tests whether the AI can find specific details buried inside a very long document (around 128k tokens — roughly a long book). Higher is better.
- Which AI model scores highest on MRCR v2 (8-needle)?
- GPT-5.5 by OpenAI holds the best MRCR v2 (8-needle) (128k average) result among tracked models, at 94.8% (released Apr 23 2026). Higher scores are better on this benchmark.
- What are the top 5 models on MRCR v2 (8-needle)?
- 1. GPT-5.5 (OpenAI) — 94.8%; 2. Claude Sonnet 4.6 (Anthropic) — 84.9%; 2. Gemini 3.1 Pro (Google) — 84.9%; 4. Gemini 3.5 Flash (Google) — 77.3%; 5. Gemini 3.0 Flash (Google) — 67.2%.
- How many models have a published MRCR v2 (8-needle) score?
- 6 tracked models have a published MRCR v2 (8-needle) (128k average) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.