Cybersecurity
ExploitGym6-hour budget
Can an AI agent turn a known software vulnerability into a working attack in a controlled lab? Built by MPI-SP researchers, the score is how many of 898 real cases (userspace programs, the V8 engine, the Linux kernel) it cracks — here with a 6-hour compute budget per case. Higher is better.
Rankings
Higher is betterExploitGym — frequently asked questions
- What is ExploitGym?
- Can an AI agent turn a known software vulnerability into a working attack in a controlled lab? Built by MPI-SP researchers, the score is how many of 898 real cases (userspace programs, the V8 engine, the Linux kernel) it cracks — here with a 6-hour compute budget per case. Higher is better.
- Which AI model scores highest on ExploitGym?
- GLM-5.3 by Z.ai holds the best ExploitGym (6-hour budget) result among tracked models, at 130 (released Aug 14 2026). Higher scores are better on this benchmark.
- What are the top 1 models on ExploitGym?
- 1. GLM-5.3 (Z.ai) — 130.
- How many models have a published ExploitGym score?
- 1 tracked model has a published ExploitGym (6-hour budget) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.