# OSWorld 2.0 — AI model rankings

Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Version 2.0 is a harder, refreshed task set. Higher is better.

5 tracked models have a published OSWorld 2.0 score. Higher is better. Scores come from published lab reports and benchmark sources; recorded sources appear beside the scores.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 5 | Anthropic | 70.6% | Lab | Jul 24 2026 |
| 2 | Muse Spark 1.3 | Meta | 66.9% | Lab | Sep 2 2026 |
| 3 | Gemini 3.8 Flash | Google | 59% | Lab | Sep 2 2026 |
| 4 | Gemini 3.7 Flash | Google | 38.1% | Lab | Aug 13 2026 |
| 5 | Qwen3.8-Flash-Next | Qwen | 19.4% | Lab | Aug 26 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/osworld-2.0
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
