Agentic computer work

Frontier-Bench v0.1

A hard, ever-evolving set of real computer tasks — coding, system administration, data work, and more — that an AI agent has to complete on its own. Run by the Harbor / Laude Institute team as the successor to Terminal-Bench (v0.1 is the first release of the task set). The score is the share of tasks solved. Higher is better.

Rankings

Higher is better

Frontier-Bench v0.1 — frequently asked questions

What is Frontier-Bench v0.1?
A hard, ever-evolving set of real computer tasks — coding, system administration, data work, and more — that an AI agent has to complete on its own. Run by the Harbor / Laude Institute team as the successor to Terminal-Bench (v0.1 is the first release of the task set). The score is the share of tasks solved. Higher is better.
Which AI model scores highest on Frontier-Bench v0.1?
Claude Opus 5 by Anthropic holds the best Frontier-Bench v0.1 result among tracked models, at 43.3% (released Jul 24 2026). Higher scores are better on this benchmark.
What are the top 5 models on Frontier-Bench v0.1?
1. Claude Opus 5 (Anthropic) — 43.3%; 2. GPT-5.6 Sol (OpenAI) — 34.4%; 3. Claude Fable 5 (Anthropic) — 33.8%; 4. Claude Opus 4.8 (Anthropic) — 21.1%; 5. GPT-5.6 Terra (OpenAI) — 20.8%.
How many models have a published Frontier-Bench v0.1 score?
9 tracked models have a published Frontier-Bench v0.1 score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks