# Gray Swan IPI (k = 1) — AI model rankings

Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. Gray Swan's indirect prompt injection benchmark measures how often such an attack succeeds when the attacker gets a single try. Lower is better.

13 tracked models have a published Gray Swan IPI (k = 1) score. Lower is better. Scores are as published at each model's release.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 5 | Anthropic | 0.2% | Lab | Jul 24 2026 |
| 2 | Claude Fable 5 | Anthropic | 0.4% | Lab | Jun 9 2026 |
| 3 | Claude Opus 4.8 | Anthropic | 0.5% | Lab | May 28 2026 |
| 4 | Claude Sonnet 5 | Anthropic | 0.6% | Lab | Jun 30 2026 |
| 5 | Muse Spark | Meta | 2.9% | Lab | Apr 8 2026 |
| 6 | GPT-5.5 | OpenAI | 3% | Lab | Apr 23 2026 |
| 7 | GPT-5.6 Sol | OpenAI | 3.1% | Lab | Jun 26 2026 |
| 8 | GPT-5.6 Terra | OpenAI | 5.4% | Lab | Jun 26 2026 |
| 9 | Gemini 3.6 Flash | Google | 7.3% | Lab | Jul 21 2026 |
| 10 | GPT-5.6 Luna | OpenAI | 8.3% | Lab | Jun 26 2026 |
| 11 | Grok 4.5 | SpaceXAI | 13.4% | Lab | Jul 8 2026 |
| 12 | Gemini 3.5 Flash | Google | 14.1% | Lab | May 19 2026 |
| 13 | Gemini 3.1 Pro | Google | 14.2% | Lab | Feb 19 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/grayswan-ipi-k1
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
