AI Model Release Tracker - Timeline of Major AI Models from 2022-2026

Claude Opus 5

Claude Opus 5 is an AI model released by Anthropic on Friday, Jul 24 2026, 24 days after Claude Sonnet 5. It has a 1M token context window. Benchmark results (shown below) cover FrontierCode v1.1 (Main), Frontier-Bench v0.1, Humanity's Last Exam, ARC-AGI-3, BioMysteryBench, OSWorld 2.0, and 6 more.

Benchmarks

Agentic coding
FrontierCode v1.1 (Main)
53.4%
Agentic computer work
Frontier-Bench v0.1
43.3%
#1
Multidisciplinary reasoning
Humanity's Last Exam
56.3%
no tools
64.7%
with tools
no tools
#1
with tools
#1
Novel problem-solving
ARC-AGI-3
30.2%
Biology
BioMysteryBench
49.4%
hard
90.1%
human solved
hard
human solved
Agentic computer use
OSWorld 2.0
70.6%
Business workflows
AutomationBench
26%
Agentic legal work
Harvey's Legal Agent Benchmark (Held-out)
11.7%
Health
HealthBench Professional
59.8%
Knowledge work
GDPval-AA v2
1861
#1
Agentic coding
DeepSWE 1.1
68.8%
#4 of 13
Web browsing
BrowseComp
90.8%
#2 of 16

About Claude Opus 5

Claude Opus 5, released July 24, 2026, brought Anthropic's Claude 5 generation to the Opus tier six weeks after Claude Fable 5 opened it — at half Fable's price, keeping Opus 4.8's $5 per million input tokens and $25 per million output. The pitch was flagship-class agentic capability at workhorse pricing: at launch it scored 43.3% on Frontier-Bench v0.1, more than double Opus 4.8's 21.1% and nearly ten points clear of Fable 5, and posted a GDPval-AA v2 Elo of 1861 for knowledge work, the best published score at the time.

The launch card leaned on breadth: 90.8% on BrowseComp for agentic search, 70.6% on OSWorld 2.0 computer use, 64.7% on Humanity's Last Exam with tools, and 30.2% on ARC-AGI-3 — which Anthropic reported as roughly three times the next best published result on the novel problem-solving benchmark. Anthropic made it the default model on Claude Max and the strongest model available on Claude Pro, positioning Opus 5 as the everyday frontier model while Fable 5 kept the edge on a handful of evaluations, including DeepSWE agentic coding and Harvey's held-out legal benchmark.

Compare Claude Opus 5 with