Next.js coding

Next.js Evals

Vercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better.

Rankings

Higher is better

Next.js Evals — frequently asked questions

What is Next.js Evals?
Vercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better.
Which AI model scores highest on Next.js Evals?
Composer 2.5 by Cursor holds the best Next.js Evals result among tracked models, at 92% (released May 18 2026). Higher scores are better on this benchmark.
What are the top 5 models on Next.js Evals?
1. Composer 2.5 (Cursor) — 92%; 1. Claude Fable 5 (Anthropic) — 92%; 1. GPT-5.6 Sol (OpenAI) — 92%; 1. Kimi K3 (Moonshot AI) — 92%; 1. Grok 4.6 (SpaceXAI) — 92%.
How many models have a published Next.js Evals score?
25 tracked models have a published Next.js Evals score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks