Safety-review circumvention

Auto-review circumvention (Internal)

OpenAI's internal safety check on how often a model finds ways around its own automated review — the guardrail that inspects what it is about to do. This one counts failures, so lower is better and zero is the goal.

Rankings

Lower is better

Auto-review circumvention (Internal) — frequently asked questions

What is Auto-review circumvention (Internal)?
OpenAI's internal safety check on how often a model finds ways around its own automated review — the guardrail that inspects what it is about to do. This one counts failures, so lower is better and zero is the goal.
Which AI model performs best on Auto-review circumvention (Internal)?
GPT-6 Astra by OpenAI holds the best Auto-review circumvention (Internal) result among tracked models, at 0% (released Sep 3 2026). Lower scores are better on this benchmark.
What are the top 1 models on Auto-review circumvention (Internal)?
1. GPT-6 Astra (OpenAI) — 0%.
How many models have a published Auto-review circumvention (Internal) score?
1 tracked model has a published Auto-review circumvention (Internal) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.