Cybersecurity
ExploitBench
A 'capability ladder' for security research, built by CMU researchers: the AI is given known bugs in Chrome's V8 engine and scored on how far it gets toward a working exploit inside a research sandbox — from understanding the patch to triggering a crash. Higher is better.
Rankings
Higher is betterExploitBench — frequently asked questions
- What is ExploitBench?
- A 'capability ladder' for security research, built by CMU researchers: the AI is given known bugs in Chrome's V8 engine and scored on how far it gets toward a working exploit inside a research sandbox — from understanding the patch to triggering a crash. Higher is better.
- Which AI model scores highest on ExploitBench?
- GLM-5.3 by Z.ai holds the best ExploitBench result among tracked models, at 54.4% (released Aug 14 2026). Higher scores are better on this benchmark.
- What are the top 1 models on ExploitBench?
- 1. GLM-5.3 (Z.ai) — 54.4%.
- How many models have a published ExploitBench score?
- 1 tracked model has a published ExploitBench score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.