Prompt injection robustness

Gray Swan IPIk = 1

Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. Gray Swan's indirect prompt injection benchmark measures how often such an attack succeeds when the attacker gets a single try. Lower is better.

Rankings

Lower is better

Gray Swan IPI — frequently asked questions

What is Gray Swan IPI?
Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. Gray Swan's indirect prompt injection benchmark measures how often such an attack succeeds when the attacker gets a single try. Lower is better.
Which AI model performs best on Gray Swan IPI?
Claude Opus 5 by Anthropic holds the best Gray Swan IPI (k = 1) result among tracked models, at 0.2% (released Jul 24 2026). Lower scores are better on this benchmark.
What are the top 5 models on Gray Swan IPI?
1. Claude Opus 5 (Anthropic) — 0.2%; 2. Claude Fable 5 (Anthropic) — 0.4%; 3. Claude Opus 4.8 (Anthropic) — 0.5%; 4. Claude Sonnet 5 (Anthropic) — 0.6%; 5. Muse Spark (Meta) — 2.9%.
How many models have a published Gray Swan IPI score?
13 tracked models have a published Gray Swan IPI (k = 1) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks