Agentic terminal coding
Terminal-Bench 4.0
Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Version 4.0 recalibrated how much time, CPU and memory each task gets, removed eight tasks and fixed nineteen, so fewer runs fail for reasons that have nothing to do with the model. Scores are not comparable with earlier versions. Higher is better.
Rankings
Higher is better#151.82%
Claude Opus 5via Terminal-Bench
244.55%
Claude Fable 5via Terminal-Bench
341.82%
GLM-5.3OWvia Terminal-Bench
437.27%
GPT-5.6 Solvia Terminal-Bench
523.64%
Claude Opus 4.8via Terminal-Bench
621.52%
GPT-5.6 Terravia Terminal-Bench
720.3%
Grok 4.6via Terminal-Bench
817.27%
GPT-5.6 Lunavia Terminal-Bench
912.42%
Claude Sonnet 5via Terminal-Bench
912.42%
Grok 4.5via Terminal-Bench
Scores marked “via” above are quoted with attribution from Terminal-Bench, retrieved 30 August 2026. Every other score on this page is the one the lab published at launch.
Terminal-Bench 4.0 — frequently asked questions
- What is Terminal-Bench 4.0?
- Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Version 4.0 recalibrated how much time, CPU and memory each task gets, removed eight tasks and fixed nineteen, so fewer runs fail for reasons that have nothing to do with the model. Scores are not comparable with earlier versions. Higher is better.
- Which AI model scores highest on Terminal-Bench 4.0?
- Claude Opus 5 by Anthropic holds the best Terminal-Bench 4.0 result among tracked models, at 51.82% (released Jul 24 2026). Higher scores are better on this benchmark.
- What are the top 5 models on Terminal-Bench 4.0?
- 1. Claude Opus 5 (Anthropic) — 51.82%; 2. Claude Fable 5 (Anthropic) — 44.55%; 3. GLM-5.3 (Z.ai) — 41.82%; 4. GPT-5.6 Sol (OpenAI) — 37.27%; 5. Claude Opus 4.8 (Anthropic) — 23.64%.
- How many models have a published Terminal-Bench 4.0 score?
- 10 tracked models have a published Terminal-Bench 4.0 score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
- What is the best open model on Terminal-Bench 4.0?
- GLM-5.3 by Z.ai is the highest-ranked model with downloadable weights on Terminal-Bench 4.0, scoring 41.82% at rank 3 overall.