Next.js coding
Next.js Evals
Vercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better.
Scores from Next.js Evals, which runs the benchmark and publishes the full field.
Rankings
Higher is better#197%
Claude Fable 5.1
#197%
GPT-6 Sol
#197%
Claude Opus 5.5
494%
Claude Opus 5
494%
Grok 4.7
690%
Gemini 3.8 Flash
690%
GPT-6 Astra
884%
Kimi K3
981%
Composer 2.5
981%
GLM-5.2
981%
Claude Sonnet 5
1277%
Claude Fable 5
1277%
GPT-5.6 Sol
1277%
GPT-6 Luna
1574%
Kimi K2.7 Code
1671%
Grok 4.6
1768%
GPT-5.3-Codex
1768%
Claude Opus 4.8
1967%
Composer 1.5
2065%
GPT-5.4
2065%
GPT-5.5-Pro
2065%
Grok 4.5
2358%
Claude Opus 4.6
2358%
Gemini 3.1 Pro
2358%
Composer 2
2358%
GLM-5.1
2358%
Claude Opus 4.7
2852%
Gemini 3.0 Pro
2852%
Kimi K2.6
3045%
Claude Sonnet 4.6
3139%
Claude Sonnet 4.5
3216%
Kimi K2.5
Next.js Evals — frequently asked questions
Vercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better.