# Frontier-Bench v0.1 — AI model rankings

A hard, ever-evolving set of real computer tasks — coding, system administration, data work, and more — that an AI agent has to complete on its own. Run by the Harbor / Laude Institute team as the successor to Terminal-Bench (v0.1 is the first release of the task set). The score is the share of tasks solved. Higher is better.

9 tracked models have a published Frontier-Bench v0.1 score. Higher is better. Scores are as published at each model's release.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 5 | Anthropic | 43.3% | Lab | Jul 24 2026 |
| 2 | GPT-5.6 Sol | OpenAI | 34.4% | Lab | Jun 26 2026 |
| 3 | Claude Fable 5 | Anthropic | 33.8% | Lab | Jun 9 2026 |
| 4 | Claude Opus 4.8 | Anthropic | 21.1% | Lab | May 28 2026 |
| 5 | GPT-5.6 Terra | OpenAI | 20.8% | Lab | Jun 26 2026 |
| 6 | Grok 4.5 | SpaceXAI | 17.8% | Lab | Jul 8 2026 |
| 7 | Claude Sonnet 5 | Anthropic | 14.6% | Lab | Jun 30 2026 |
| 8 | GPT-5.6 Luna | OpenAI | 14.3% | Lab | Jun 26 2026 |
| 9 | GLM-5.2 | Z.ai | 5.1% | Lab | Jun 16 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/frontier-bench-v0.1
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
