Nonsense detection
BullshitBench v2
Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
Rankings
Higher is better#195%
Claude Opus 4.8
291%
Claude Sonnet 4.6
390%
Claude Opus 4.5
487%
Claude Opus 4.6
583%
Claude Opus 4.7
680%
Claude Sonnet 5
779%
Claude Sonnet 4.5
877%
Claude Haiku 4.5
973%
Kimi K3
973%
Claude Opus 5
1172%
Qwen3.6-Plus
1271%
Qwen3.7-Max
1365%
Kimi K2.6
1365%
Gemini 3.5 Flash-Lite
1556%
Grok 4.20 Beta
1654%
Claude Fable 5
1654%
Grok 4.5
1853%
GPT-5.6 Terra
1952%
Kimi K2.5
2050%
Claude 3.5 Haiku
2050%
Grok 4.3 Beta
2249%
Claude 3.7 Sonnet
2348%
GPT-5.4
2447%
GPT-5.5
2447%
GPT-5.6 Sol
2645%
Claude 3.5 Sonnet
2743%
Claude Opus 4.1
2840%
GPT-5.6 Luna
2939%
GPT-5-Codex
2939%
Gemini 3.6 Flash
3138%
GPT-5.2
3138%
DeepSeek-V4-Flash-0731
3337%
Gemini 3.1 Pro
3436%
GPT-5.5-Pro
3534%
Claude Opus 4
3632%
GPT-5.4 mini
3731%
GLM-5.2
3830%
Claude Sonnet 4
3928%
LLaMA 4 Maverick
3928%
GLM-5
4126%
o3
4225%
GPT-5.1
4324%
GPT-5.3-Codex
4422%
GLM-5.1
4521%
GPT-5
4620%
Gemini 2.5 Pro
4620%
Qwen3-CoderOW
4620%
Gemini 3.5 Flash
4919%
LLaMA 4 Scout
4919%
Gemini 2.5 Flash
4919%
Grok 4.1 Fast
5218%
DeepSeek-V4-Flash
5315%
Gemini 2.0 Flash
5414%
LLaMA 3.1
5414%
GPT-4.1
5414%
GPT-5.4 nano
5414%
DeepSeek-V4-Pro
5813%
DeepSeek-V3.2
5912%
GPT-4o
6011%
gpt-oss-120b
6011%
Gemini 3.1 Flash-Lite
6210%
Claude 3 Haiku
6210%
Kimi K2
648%
o4-mini
648%
DeepSeek R1-0528
648%
GLM-4.5
672%
GPT-4o mini
BullshitBench v2 — frequently asked questions
- What is BullshitBench v2?
- Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
- Which AI model scores highest on BullshitBench v2?
- Claude Opus 4.8 by Anthropic holds the best BullshitBench v2 result among tracked models, at 95% (released May 28 2026). Higher scores are better on this benchmark.
- What are the top 5 models on BullshitBench v2?
- 1. Claude Opus 4.8 (Anthropic) — 95%; 2. Claude Sonnet 4.6 (Anthropic) — 91%; 3. Claude Opus 4.5 (Anthropic) — 90%; 4. Claude Opus 4.6 (Anthropic) — 87%; 5. Claude Opus 4.7 (Anthropic) — 83%.
- How many models have a published BullshitBench v2 score?
- 67 tracked models have a published BullshitBench v2 score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
- What is the best open model on BullshitBench v2?
- Qwen3-Coder by Qwen is the highest-ranked model with downloadable weights on BullshitBench v2, scoring 20% at rank 46 overall.