# AutomationBench — AI model rankings

Tests whether the AI can run real multi-step business workflows — the kind of end-to-end office processes companies want to automate — from start to finish. Higher is better.

10 tracked models have a published AutomationBench score. Higher is better. Scores are as published at each model's release.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | Qwen3.8-Max-0902 | Qwen | 50.8% | Lab | Sep 2 2026 |
| 2 | Muse Spark 1.3 | Meta | 49.4% | Lab | Sep 2 2026 |
| 3 | GLM-5.3-Flash | Z.ai | 48.8% | Lab | Aug 26 2026 |
| 4 | GLM-5.3 | Z.ai | 48.2% | Lab | Aug 14 2026 |
| 5 | GPT-6 Astra | OpenAI | 41.4% | Lab | Sep 3 2026 |
| 6 | DeepSeek-V4-Pro-0813 | DeepSeek | 31.8% | Lab | Aug 13 2026 |
| 7 | Claude Fable 5.1 | Anthropic | 31.4% | Lab | Sep 1 2026 |
| 8 | Gemini 3.7 Flash | Google | 30.4% | Lab | Aug 13 2026 |
| 9 | Claude Opus 5 | Anthropic | 26% | Lab | Jul 24 2026 |
| 10 | DeepSeek-V4-Flash-0731 | DeepSeek | 25.1% | Lab | Jul 31 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/automationbench
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
