Compare AI models

Specifications
Context window
1M—
Benchmarks
SWE-Bench Verified
95.5%77.4%
Terminal-Bench 2.1
88%67.3%
Humanity's Last Exam · with tools
64.5%50.4%
OSWorld-Verified
85%53.3%
Overview
CompanyAnthropicMeta
Release dateJun 9 2026Apr 8 2026
AccessClosedClosed
Model detailsView modelView model

Frequently asked questions

Claude Mythos 5 leads Muse Spark on 4 of the 4 benchmarks they both report (SWE-Bench Verified, Terminal-Bench 2.1, Humanity's Last Exam, OSWorld-Verified). Muse Spark shipped 62 days before Claude Mythos 5, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On SWE-Bench Verified, Claude Mythos 5 leads at 95.5% vs Muse Spark at 77.4%. On Terminal-Bench 2.1, Claude Mythos 5 leads at 88% vs Muse Spark at 67.3%. On Humanity's Last Exam · with tools, Claude Mythos 5 leads at 64.5% vs Muse Spark at 50.4%. On OSWorld-Verified, Claude Mythos 5 leads at 85% vs Muse Spark at 53.3%.