Nonsense detection

BullshitBench v2

Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.

Rankings

Higher is better

BullshitBench v2 — frequently asked questions

What is BullshitBench v2?
Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
Which AI model scores highest on BullshitBench v2?
Claude Opus 4.8 by Anthropic holds the best BullshitBench v2 result among tracked models, at 95% (released May 28 2026). Higher scores are better on this benchmark.
What are the top 5 models on BullshitBench v2?
1. Claude Opus 4.8 (Anthropic) — 95%; 2. Claude Sonnet 4.6 (Anthropic) — 91%; 3. Claude Opus 4.5 (Anthropic) — 90%; 4. Claude Opus 4.6 (Anthropic) — 87%; 5. Claude Opus 4.7 (Anthropic) — 83%.
How many models have a published BullshitBench v2 score?
67 tracked models have a published BullshitBench v2 score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
What is the best open model on BullshitBench v2?
Qwen3-Coder by Qwen is the highest-ranked model with downloadable weights on BullshitBench v2, scoring 20% at rank 46 overall.
← All benchmarks