Safety-review circumvention
Auto-review circumvention (Internal)
OpenAI's internal safety check on how often a model finds ways around its own automated review — the guardrail that inspects what it is about to do. This one counts failures, so lower is better and zero is the goal.
Rankings
Lower is betterAuto-review circumvention (Internal) — frequently asked questions
- What is Auto-review circumvention (Internal)?
- OpenAI's internal safety check on how often a model finds ways around its own automated review — the guardrail that inspects what it is about to do. This one counts failures, so lower is better and zero is the goal.
- Which AI model performs best on Auto-review circumvention (Internal)?
- GPT-6 Astra by OpenAI holds the best Auto-review circumvention (Internal) result among tracked models, at 0% (released Sep 3 2026). Lower scores are better on this benchmark.
- What are the top 1 models on Auto-review circumvention (Internal)?
- 1. GPT-6 Astra (OpenAI) — 0%.
- How many models have a published Auto-review circumvention (Internal) score?
- 1 tracked model has a published Auto-review circumvention (Internal) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.