Terminal-Bench 4.0
Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Version 4.0 recalibrated how much time, CPU and memory each task gets, removed eight tasks and fixed nineteen, so fewer runs fail for reasons that have nothing to do with the model. Scores are not comparable with earlier versions. Higher is better.
Rankings
Higher is betterScores marked “via” above are quoted with attribution from Terminal-Bench, retrieved 5 September 2026. Every other score on this page is the one the lab published at launch.
Terminal-Bench 4.0 — frequently asked questions
Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Version 4.0 recalibrated how much time, CPU and memory each task gets, removed eight tasks and fixed nineteen, so fewer runs fail for reasons that have nothing to do with the model. Scores are not comparable with earlier versions. Higher is better.