Cybersecurity

ExploitGym2-hour budget

Can an AI agent turn a known software vulnerability into a working attack in a controlled lab? Built by MPI-SP researchers, the score is how many of 898 real cases (userspace programs, the V8 engine, the Linux kernel) it cracks — here with a 2-hour compute budget per case. Higher is better.

Rankings

Higher is better

ExploitGym — frequently asked questions

What is ExploitGym?
Can an AI agent turn a known software vulnerability into a working attack in a controlled lab? Built by MPI-SP researchers, the score is how many of 898 real cases (userspace programs, the V8 engine, the Linux kernel) it cracks — here with a 2-hour compute budget per case. Higher is better.
Which AI model scores highest on ExploitGym?
GLM-5.3 by Z.ai holds the best ExploitGym (2-hour budget) result among tracked models, at 105 (released Aug 14 2026). Higher scores are better on this benchmark.
What are the top 1 models on ExploitGym?
1. GLM-5.3 (Z.ai) — 105.
How many models have a published ExploitGym score?
1 tracked model has a published ExploitGym (2-hour budget) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks