AI Model Release Tracker - Timeline of Major AI Models from 2022-2026

Agentic computer work

Frontier-Bench v0.1

A hard, ever-evolving set of real computer tasks — coding, system administration, data work, and more — that an AI agent has to complete on its own. Run by the Harbor / Laude Institute team as the successor to Terminal-Bench (v0.1 is the first release of the task set). The score is the share of tasks solved. Higher is better.

Rankings

Higher is better
← All benchmarks