Nonsense detection
BullshitBench v2
Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
Scores from BullshitBench, which runs the benchmark and publishes the full field.
Rankings
Higher is better#195%
Claude Opus 4.8
#195%
Qwen3.8-Max
391%
Claude Sonnet 4.6
490%
Claude Opus 4.5
587%
Claude Opus 4.6
683%
Claude Opus 4.7
780%
Claude Sonnet 5
879%
Claude Sonnet 4.5
879%
Qwen3.8-27BOW
1077%
Claude Haiku 4.5
1077%
Claude Fable 5.1
1274%
Kimi K3
1373%
Claude Opus 5
1472%
Qwen3.6-Plus
1472%
Qwen3.7-Max
1472%
GLM-5.3OW
1769%
GPT-6 Astra
1866%
Gemini 3.5 Flash-Lite
1866%
Grok 4.6
2065%
Kimi K2.6
2157%
DeepSeek-V4.1-Flash
2256%
Grok 4.20 Beta
2256%
Claude Fable 5
2455%
Grok 4.5
2554%
GPT-5.6 Terra
2652%
Kimi K2.5
2750%
Claude 3.5 Haiku
2750%
Grok 4.3 Beta
2750%
Muse Spark 1.2
3049%
Claude 3.7 Sonnet
3148%
Gemini 3.0 Pro
3148%
GPT-5.4
3148%
GPT-5.6 Sol
3447%
GPT-5.5
3546%
Claude 3.5 Sonnet
3643%
Claude Opus 4.1
3741%
GPT-5.6 Luna
3839%
GPT-5-Codex
3839%
Gemini 3.6 Flash
3839%
DeepSeek-V4-Flash-0731
4138%
GPT-5.2
4237%
Gemini 3.1 Pro
4336%
GPT-5.5-Pro
4435%
Gemini 3.7 Flash
4435%
DeepSeek-V4-Pro-0813
4634%
Claude Opus 4
4732%
GPT-5.4 mini
4831%
GLM-5.2
4930%
Claude Sonnet 4
5029%
LLaMA 4 Maverick
5128%
GLM-5
5226%
o3
5325%
GPT-5.1
5424%
GPT-5.3-Codex
5522%
GLM-5.1
5621%
GPT-5
5720%
Gemini 2.5 Pro
5720%
LLaMA 4 Scout
5720%
Qwen3-CoderOW
5720%
Gemini 3.5 Flash
6119%
Gemini 2.5 Flash
6119%
Grok 4.1 Fast
6318%
DeepSeek-V4-Flash
6415%
Gemini 2.0 Flash
6514%
LLaMA 3.1
6514%
GPT-4.1
6514%
GPT-5.4 nano
6514%
DeepSeek-V4-Pro
6913%
DeepSeek-V3.2
7012%
GPT-4o
7012%
gpt-oss-120bOW
7211%
Gemini 3.1 Flash-Lite
7310%
Claude 3 Haiku
7310%
Kimi K2
7310%
Gemini 3.0 Flash
768%
o4-mini
768%
DeepSeek R1-0528
768%
GLM-4.5
792%
GPT-4o mini
BullshitBench v2 — frequently asked questions
Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.