# CyberGym — AI model rankings

Tests the AI on cybersecurity challenges — finding and exploiting software weaknesses inside a safe sandbox. Higher is better.

4 tracked models have a published CyberGym score. Higher is better. Scores are as published at each model's release.

## Ranking

| Rank | Model | Developer | Score | Released |
| --- | --- | --- | --- | --- |
| 1 | GPT-5.5 | OpenAI | 81.8% | Apr 23 2026 |
| 2 | GPT-5.4 | OpenAI | 79% | Mar 5 2026 |
| 3 | DeepSeek-V4-Flash-0731 | DeepSeek | 76.7% | Jul 31 2026 |
| 4 | Claude Opus 4.7 | Anthropic | 73.1% | Apr 16 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/cybergym
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
