# SRE-Bench — AI model rankings

Can the AI keep production running? It is dropped into a broken Kubernetes system and has to diagnose the incident and fix it safely, the way an on-call site-reliability engineer would. Scored here on the best of four attempts. Higher is better.

1 tracked model have a published SRE-Bench score. Higher is better. Scores are as published at each model's release.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | GPT-6 Astra | OpenAI | 99.2% | Lab | Sep 3 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/sre-bench
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
