# ProgramBench — AI model rankings

The AI receives a working program and its documentation, then builds a replacement from scratch without the original source code, internet access or decompilation. The score is the percentage of 200 programs that pass every behavioral test. We record each model's best published mini-SWE-agent result, including higher reasoning efforts where available. Partial test-pass rates and almost-solved programs do not count toward this score. Equal scores share a rank here; the official board also uses partial progress to break ties. Higher is better.

Scores from [ProgramBench](https://programbench.com/), which runs the benchmark and publishes the full field.

17 tracked models have a published ProgramBench score. Higher is better. Scores are gathered from ProgramBench; recorded retrieval dates appear beside the scores.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 5 | Anthropic | 4.5% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Jul 24 2026 |
| 2 | GPT-5.6 Sol | OpenAI | 1% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Jun 26 2026 |
| 3 | GPT-5.5 | OpenAI | 0.5% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Apr 23 2026 |
| 3 | Gemini 3.6 Flash | Google | 0.5% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Jul 21 2026 |
| 5 | GPT-5 mini | OpenAI | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Aug 7 2025 |
| 5 | Claude Haiku 4.5 | Anthropic | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Oct 15 2025 |
| 5 | Gemini 3.0 Flash | Google | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Dec 17 2025 |
| 5 | Claude Opus 4.6 | Anthropic | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Feb 5 2026 |
| 5 | Claude Sonnet 4.6 | Anthropic | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Feb 17 2026 |
| 5 | Gemini 3.1 Pro | Google | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Feb 19 2026 |
| 5 | GPT-5.4 | OpenAI | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Mar 5 2026 |
| 5 | GPT-5.4 mini | OpenAI | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Mar 17 2026 |
| 5 | Claude Opus 4.7 | Anthropic | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Apr 16 2026 |
| 5 | Gemini 3.5 Flash | Google | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | May 19 2026 |
| 5 | Claude Opus 4.8 | Anthropic | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | May 28 2026 |
| 5 | GLM-5.2 | Z.ai | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Jun 16 2026 |
| 5 | Gemini 3.7 Flash | Google | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | Aug 13 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/programbench
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
