# Terminal-Bench-Science 0.1 — AI model rankings

The same command-line setup as Terminal-Bench, pointed at scientific work: the AI has to drive research tooling and computational workflows through to a result, rather than administer a machine. Version 0.1 is the first release of the task set, and scores run lower than on the general board. Higher is better.

3 tracked models have a published Terminal-Bench-Science 0.1 score. Higher is better. Scores are as published at each model's release.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | Claude Fable 5.1 | Anthropic | 52.6% | Lab | Sep 1 2026 |
| 2 | Claude Opus 5 | Anthropic | 29% | Lab | Jul 24 2026 |
| 3 | Claude Fable 5 | Anthropic | 24.7% | Lab | Jun 9 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/terminal-bench-science-0.1
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
