Long context

MRCR v2 (8-needle)1M pointwise

Tests whether the AI can find specific details buried inside an enormous document (around 1 million tokens — many books). Higher is better.

Rankings

Higher is better

MRCR v2 (8-needle) — frequently asked questions

What is MRCR v2 (8-needle)?
Tests whether the AI can find specific details buried inside an enormous document (around 1 million tokens — many books). Higher is better.
Which AI model scores highest on MRCR v2 (8-needle)?
Gemini 3.5 Flash by Google holds the best MRCR v2 (8-needle) (1M pointwise) result among tracked models, at 26.6% (released May 19 2026). Higher scores are better on this benchmark.
What are the top 3 models on MRCR v2 (8-needle)?
1. Gemini 3.5 Flash (Google) — 26.6%; 2. Gemini 3.1 Pro (Google) — 26.3%; 3. Gemini 3.0 Flash (Google) — 22.1%.
How many models have a published MRCR v2 (8-needle) score?
3 tracked models have a published MRCR v2 (8-needle) (1M pointwise) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks