Prompt injection robustness

Gray Swan IPIk = 15

Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. This variant gives the attacker 15 tries and counts an attack as successful if any of them works. Lower is better.

Rankings

Lower is better

Gray Swan IPI — frequently asked questions

What is Gray Swan IPI?
Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. This variant gives the attacker 15 tries and counts an attack as successful if any of them works. Lower is better.
Which AI model performs best on Gray Swan IPI?
Claude Opus 5 by Anthropic holds the best Gray Swan IPI (k = 15) result among tracked models, at 2% (released Jul 24 2026). Lower scores are better on this benchmark.
What are the top 5 models on Gray Swan IPI?
1. Claude Opus 5 (Anthropic) — 2%; 2. Claude Fable 5 (Anthropic) — 2.8%; 3. Claude Opus 4.8 (Anthropic) — 5.5%; 4. Claude Sonnet 5 (Anthropic) — 5.9%; 5. Muse Spark (Meta) — 16.5%.
How many models have a published Gray Swan IPI score?
13 tracked models have a published Gray Swan IPI (k = 15) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks