Repo-level code generation

NL2Repo-Bench

Tests whether the AI can turn a natural-language requirement into working code across an entire repository, not just produce a single function or patch. Higher is better.

Rankings

Higher is better

NL2Repo-Bench — frequently asked questions

What is NL2Repo-Bench?
Tests whether the AI can turn a natural-language requirement into working code across an entire repository, not just produce a single function or patch. Higher is better.
Which AI model scores highest on NL2Repo-Bench?
Qwen3.8-27B by Qwen holds the best NL2Repo-Bench result among tracked models, at 42.3% (released Aug 14 2026). Higher scores are better on this benchmark.
What are the top 1 models on NL2Repo-Bench?
1. Qwen3.8-27B (Qwen) — 42.3%.
How many models have a published NL2Repo-Bench score?
1 tracked model has a published NL2Repo-Bench score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
What is the best open model on NL2Repo-Bench?
Qwen3.8-27B by Qwen is the highest-ranked model with downloadable weights on NL2Repo-Bench, scoring 42.3% at rank 1 overall.
← All benchmarks