Repo-level code generation
NL2Repo-Bench
Tests whether the AI can turn a natural-language requirement into working code across an entire repository, not just produce a single function or patch. Higher is better.
Rankings
Higher is betterNL2Repo-Bench — frequently asked questions
- What is NL2Repo-Bench?
- Tests whether the AI can turn a natural-language requirement into working code across an entire repository, not just produce a single function or patch. Higher is better.
- Which AI model scores highest on NL2Repo-Bench?
- Qwen3.8-27B by Qwen holds the best NL2Repo-Bench result among tracked models, at 42.3% (released Aug 14 2026). Higher scores are better on this benchmark.
- What are the top 1 models on NL2Repo-Bench?
- 1. Qwen3.8-27B (Qwen) — 42.3%.
- How many models have a published NL2Repo-Bench score?
- 1 tracked model has a published NL2Repo-Bench score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
- What is the best open model on NL2Repo-Bench?
- Qwen3.8-27B by Qwen is the highest-ranked model with downloadable weights on NL2Repo-Bench, scoring 42.3% at rank 1 overall.