Next.js coding
Next.js Evals
Vercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better.
Rankings
Higher is better#192%
Composer 2.5
#192%
Claude Fable 5
#192%
GPT-5.6 Sol
#192%
Kimi K3
#192%
Grok 4.6
688%
Claude Opus 4.8
688%
GLM-5.2
688%
Claude Opus 5
983%
GPT-5.3-Codex
983%
GPT-5.4
983%
GPT-5.5-Pro
983%
Grok 4.5
1379%
Claude Sonnet 5
1475%
Claude Opus 4.6
1475%
Gemini 3.1 Pro
1475%
Composer 2
1475%
GLM-5.1
1475%
Claude Opus 4.7
1475%
Kimi K2.7 Code
2067%
Gemini 3.0 Pro
2067%
Composer 1.5
2067%
Kimi K2.6
2358%
Claude Sonnet 4.6
2450%
Claude Sonnet 4.5
2521%
Kimi K2.5
Next.js Evals — frequently asked questions
- What is Next.js Evals?
- Vercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better.
- Which AI model scores highest on Next.js Evals?
- Composer 2.5 by Cursor holds the best Next.js Evals result among tracked models, at 92% (released May 18 2026). Higher scores are better on this benchmark.
- What are the top 5 models on Next.js Evals?
- 1. Composer 2.5 (Cursor) — 92%; 1. Claude Fable 5 (Anthropic) — 92%; 1. GPT-5.6 Sol (OpenAI) — 92%; 1. Kimi K3 (Moonshot AI) — 92%; 1. Grok 4.6 (SpaceXAI) — 92%.
- How many models have a published Next.js Evals score?
- 25 tracked models have a published Next.js Evals score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.