Site reliability

SRE-Bench

Can the AI keep production running? It is dropped into a broken Kubernetes system and has to diagnose the incident and fix it safely, the way an on-call site-reliability engineer would. Scored here on the best of four attempts. Higher is better.

Rankings

Higher is better

SRE-Bench — frequently asked questions

What is SRE-Bench?
Can the AI keep production running? It is dropped into a broken Kubernetes system and has to diagnose the incident and fix it safely, the way an on-call site-reliability engineer would. Scored here on the best of four attempts. Higher is better.
Which AI model scores highest on SRE-Bench?
GPT-6 Astra by OpenAI holds the best SRE-Bench result among tracked models, at 99.2% (released Sep 3 2026). Higher scores are better on this benchmark.
What are the top 1 models on SRE-Bench?
1. GPT-6 Astra (OpenAI) — 99.2%.
How many models have a published SRE-Bench score?
1 tracked model has a published SRE-Bench score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.