# Gray Swan IPI (k = 15) — AI model rankings

Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. This variant gives the attacker 15 tries and counts an attack as successful if any of them works. Lower is better.

13 tracked models have a published Gray Swan IPI (k = 15) score. Lower is better. Scores are as published at each model's release.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 5 | Anthropic | 2% | Lab | Jul 24 2026 |
| 2 | Claude Fable 5 | Anthropic | 2.8% | Lab | Jun 9 2026 |
| 3 | Claude Opus 4.8 | Anthropic | 5.5% | Lab | May 28 2026 |
| 4 | Claude Sonnet 5 | Anthropic | 5.9% | Lab | Jun 30 2026 |
| 5 | Muse Spark | Meta | 16.5% | Lab | Apr 8 2026 |
| 6 | GPT-5.6 Sol | OpenAI | 20% | Lab | Jun 26 2026 |
| 7 | GPT-5.5 | OpenAI | 20.8% | Lab | Apr 23 2026 |
| 8 | GPT-5.6 Terra | OpenAI | 30.4% | Lab | Jun 26 2026 |
| 9 | Gemini 3.6 Flash | Google | 37.3% | Lab | Jul 21 2026 |
| 10 | GPT-5.6 Luna | OpenAI | 43.9% | Lab | Jun 26 2026 |
| 11 | Gemini 3.1 Pro | Google | 49.2% | Lab | Feb 19 2026 |
| 12 | Gemini 3.5 Flash | Google | 60.5% | Lab | May 19 2026 |
| 13 | Grok 4.5 | SpaceXAI | 60.8% | Lab | Jul 8 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/grayswan-ipi-k15
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
