Agentic coding
FrontierCode v1.1 (Main)
A set of very hard, frontier-difficulty coding tasks an AI agent has to complete end to end. The score is the share of tasks in the main split it solves. Higher is better.
Rankings
Higher is betterFrontierCode v1.1 (Main) — frequently asked questions
- What is FrontierCode v1.1 (Main)?
- A set of very hard, frontier-difficulty coding tasks an AI agent has to complete end to end. The score is the share of tasks in the main split it solves. Higher is better.
- Which AI model scores highest on FrontierCode v1.1 (Main)?
- Claude Opus 5 by Anthropic holds the best FrontierCode v1.1 (Main) result among tracked models, at 53.4% (released Jul 24 2026). Higher scores are better on this benchmark.
- What are the top 1 models on FrontierCode v1.1 (Main)?
- 1. Claude Opus 5 (Anthropic) — 53.4%.
- How many models have a published FrontierCode v1.1 (Main) score?
- 1 tracked model has a published FrontierCode v1.1 (Main) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.