Cybersecurity

ExploitBench

A 'capability ladder' for security research, built by CMU researchers: the AI is given known bugs in Chrome's V8 engine and scored on how far it gets toward a working exploit inside a research sandbox — from understanding the patch to triggering a crash. Higher is better.

Rankings

Higher is better

ExploitBench — frequently asked questions

What is ExploitBench?
A 'capability ladder' for security research, built by CMU researchers: the AI is given known bugs in Chrome's V8 engine and scored on how far it gets toward a working exploit inside a research sandbox — from understanding the patch to triggering a crash. Higher is better.
Which AI model scores highest on ExploitBench?
GLM-5.3 by Z.ai holds the best ExploitBench result among tracked models, at 54.4% (released Aug 14 2026). Higher scores are better on this benchmark.
What are the top 1 models on ExploitBench?
1. GLM-5.3 (Z.ai) — 54.4%.
How many models have a published ExploitBench score?
1 tracked model has a published ExploitBench score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks