Cybersecurity

ExploitGym6-hour budget

Can an AI agent turn a known software vulnerability into a working attack in a controlled lab? Built by MPI-SP researchers, the score is how many of 898 real cases (userspace programs, the V8 engine, the Linux kernel) it cracks — here with a 6-hour compute budget per case. Higher is better.

Rankings

Higher is better

ExploitGym — frequently asked questions

What is ExploitGym?
Can an AI agent turn a known software vulnerability into a working attack in a controlled lab? Built by MPI-SP researchers, the score is how many of 898 real cases (userspace programs, the V8 engine, the Linux kernel) it cracks — here with a 6-hour compute budget per case. Higher is better.
Which AI model scores highest on ExploitGym?
GLM-5.3 by Z.ai holds the best ExploitGym (6-hour budget) result among tracked models, at 130 (released Aug 14 2026). Higher scores are better on this benchmark.
What are the top 1 models on ExploitGym?
1. GLM-5.3 (Z.ai) — 130.
How many models have a published ExploitGym score?
1 tracked model has a published ExploitGym (6-hour budget) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks