# Terminal-Bench 4.0 — AI model rankings

Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Version 4.0 recalibrated how much time, CPU and memory each task gets, removed eight tasks and fixed nineteen, so fewer runs fail for reasons that have nothing to do with the model. Scores are not comparable with earlier versions. Higher is better.

10 tracked models have a published Terminal-Bench 4.0 score. Higher is better. Scores are as published at each model's release.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 5 | Anthropic | 51.82% | [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), retrieved 2026-08-30 | Jul 24 2026 |
| 2 | Claude Fable 5 | Anthropic | 44.55% | [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), retrieved 2026-08-30 | Jun 9 2026 |
| 3 | GLM-5.3 | Z.ai | 41.82% | [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), retrieved 2026-08-30 | Aug 14 2026 |
| 4 | GPT-5.6 Sol | OpenAI | 37.27% | [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), retrieved 2026-08-30 | Jun 26 2026 |
| 5 | Claude Opus 4.8 | Anthropic | 23.64% | [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), retrieved 2026-08-30 | May 28 2026 |
| 6 | GPT-5.6 Terra | OpenAI | 21.52% | [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), retrieved 2026-08-30 | Jun 26 2026 |
| 7 | Grok 4.6 | SpaceXAI | 20.3% | [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), retrieved 2026-08-30 | Aug 12 2026 |
| 8 | GPT-5.6 Luna | OpenAI | 17.27% | [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), retrieved 2026-08-30 | Jun 26 2026 |
| 9 | Claude Sonnet 5 | Anthropic | 12.42% | [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), retrieved 2026-08-30 | Jun 30 2026 |
| 9 | Grok 4.5 | SpaceXAI | 12.42% | [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), retrieved 2026-08-30 | Jul 8 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/terminal-bench-4.0
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
