Agentic scientific computing
Terminal-Bench-Science 0.1
The same command-line setup as Terminal-Bench, pointed at scientific work: the AI has to drive research tooling and computational workflows through to a result, rather than administer a machine. Version 0.1 is the first release of the task set, and scores run lower than on the general board. Higher is better.
Rankings
Higher is betterTerminal-Bench-Science 0.1 — frequently asked questions
- What is Terminal-Bench-Science 0.1?
- The same command-line setup as Terminal-Bench, pointed at scientific work: the AI has to drive research tooling and computational workflows through to a result, rather than administer a machine. Version 0.1 is the first release of the task set, and scores run lower than on the general board. Higher is better.
- Which AI model scores highest on Terminal-Bench-Science 0.1?
- Claude Fable 5.1 by Anthropic holds the best Terminal-Bench-Science 0.1 result among tracked models, at 52.6% (released Sep 1 2026). Higher scores are better on this benchmark.
- What are the top 3 models on Terminal-Bench-Science 0.1?
- 1. Claude Fable 5.1 (Anthropic) — 52.6%; 2. Claude Opus 5 (Anthropic) — 29%; 3. Claude Fable 5 (Anthropic) — 24.7%.
- How many models have a published Terminal-Bench-Science 0.1 score?
- 3 tracked models have a published Terminal-Bench-Science 0.1 score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.