# BullshitBench v2 — AI model rankings

Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.

Scores from [BullshitBench](https://github.com/petergpt/bullshit-benchmark), which runs the benchmark and publishes the full field.

79 tracked models have a published BullshitBench v2 score. Higher is better. Scores are gathered from BullshitBench; recorded retrieval dates appear beside the scores.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 4.8 | Anthropic | 95% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | May 28 2026 |
| 1 | Qwen3.8-Max | Qwen | 95% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Aug 3 2026 |
| 3 | Claude Sonnet 4.6 | Anthropic | 91% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Feb 17 2026 |
| 4 | Claude Opus 4.5 | Anthropic | 90% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Nov 24 2025 |
| 5 | Claude Opus 4.6 | Anthropic | 87% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Feb 5 2026 |
| 6 | Claude Opus 4.7 | Anthropic | 83% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 16 2026 |
| 7 | Claude Sonnet 5 | Anthropic | 80% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jun 30 2026 |
| 8 | Claude Sonnet 4.5 | Anthropic | 79% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Sep 29 2025 |
| 8 | Qwen3.8-27B | Qwen | 79% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Aug 14 2026 |
| 10 | Claude Haiku 4.5 | Anthropic | 77% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Oct 15 2025 |
| 10 | Claude Fable 5.1 | Anthropic | 77% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Sep 1 2026 |
| 12 | Kimi K3 | Moonshot AI | 74% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jul 16 2026 |
| 13 | Claude Opus 5 | Anthropic | 73% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jul 24 2026 |
| 14 | Qwen3.6-Plus | Qwen | 72% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 2 2026 |
| 14 | Qwen3.7-Max | Qwen | 72% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | May 20 2026 |
| 14 | GLM-5.3 | Z.ai | 72% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Aug 14 2026 |
| 17 | GPT-6 Astra | OpenAI | 69% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Sep 3 2026 |
| 18 | Gemini 3.5 Flash-Lite | Google | 66% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jul 21 2026 |
| 18 | Grok 4.6 | SpaceXAI | 66% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Aug 12 2026 |
| 20 | Kimi K2.6 | Moonshot AI | 65% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 21 2026 |
| 21 | DeepSeek-V4.1-Flash | DeepSeek | 57% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Sep 10 2026 |
| 22 | Grok 4.20 Beta | SpaceXAI | 56% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Feb 17 2026 |
| 22 | Claude Fable 5 | Anthropic | 56% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jun 9 2026 |
| 24 | Grok 4.5 | SpaceXAI | 55% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jul 8 2026 |
| 25 | GPT-5.6 Terra | OpenAI | 54% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jun 26 2026 |
| 26 | Kimi K2.5 | Moonshot AI | 52% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jan 27 2026 |
| 27 | Claude 3.5 Haiku | Anthropic | 50% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Oct 22 2024 |
| 27 | Grok 4.3 Beta | SpaceXAI | 50% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 17 2026 |
| 27 | Muse Spark 1.2 | Meta | 50% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Aug 5 2026 |
| 30 | Claude 3.7 Sonnet | Anthropic | 49% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Feb 24 2025 |
| 31 | Gemini 3.0 Pro | Google | 48% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Nov 18 2025 |
| 31 | GPT-5.4 | OpenAI | 48% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Mar 5 2026 |
| 31 | GPT-5.6 Sol | OpenAI | 48% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jun 26 2026 |
| 34 | GPT-5.5 | OpenAI | 47% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 23 2026 |
| 35 | Claude 3.5 Sonnet | Anthropic | 46% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jun 20 2024 |
| 36 | Claude Opus 4.1 | Anthropic | 43% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Aug 5 2025 |
| 37 | GPT-5.6 Luna | OpenAI | 41% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jun 26 2026 |
| 38 | GPT-5-Codex | OpenAI | 39% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Sep 15 2025 |
| 38 | Gemini 3.6 Flash | Google | 39% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jul 21 2026 |
| 38 | DeepSeek-V4-Flash-0731 | DeepSeek | 39% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jul 31 2026 |
| 41 | GPT-5.2 | OpenAI | 38% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Dec 11 2025 |
| 42 | Gemini 3.1 Pro | Google | 37% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Feb 19 2026 |
| 43 | GPT-5.5-Pro | OpenAI | 36% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 23 2026 |
| 44 | Gemini 3.7 Flash | Google | 35% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Aug 13 2026 |
| 44 | DeepSeek-V4-Pro-0813 | DeepSeek | 35% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Aug 13 2026 |
| 46 | Claude Opus 4 | Anthropic | 34% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | May 22 2025 |
| 47 | GPT-5.4 mini | OpenAI | 32% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Mar 17 2026 |
| 48 | GLM-5.2 | Z.ai | 31% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jun 16 2026 |
| 49 | Claude Sonnet 4 | Anthropic | 30% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | May 22 2025 |
| 50 | LLaMA 4 Maverick | Meta | 29% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 5 2025 |
| 51 | GLM-5 | Z.ai | 28% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Feb 12 2026 |
| 52 | o3 | OpenAI | 26% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 16 2025 |
| 53 | GPT-5.1 | OpenAI | 25% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Nov 12 2025 |
| 54 | GPT-5.3-Codex | OpenAI | 24% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Feb 5 2026 |
| 55 | GLM-5.1 | Z.ai | 22% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 7 2026 |
| 56 | GPT-5 | OpenAI | 21% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Aug 7 2025 |
| 57 | Gemini 2.5 Pro | Google | 20% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Mar 25 2025 |
| 57 | LLaMA 4 Scout | Meta | 20% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 5 2025 |
| 57 | Qwen3-Coder | Qwen | 20% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jul 22 2025 |
| 57 | Gemini 3.5 Flash | Google | 20% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | May 19 2026 |
| 61 | Gemini 2.5 Flash | Google | 19% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 17 2025 |
| 61 | Grok 4.1 Fast | SpaceXAI | 19% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Nov 19 2025 |
| 63 | DeepSeek-V4-Flash | DeepSeek | 18% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 24 2026 |
| 64 | Gemini 2.0 Flash | Google | 15% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jan 30 2025 |
| 65 | LLaMA 3.1 | Meta | 14% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jul 23 2024 |
| 65 | GPT-4.1 | OpenAI | 14% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 14 2025 |
| 65 | GPT-5.4 nano | OpenAI | 14% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Mar 17 2026 |
| 65 | DeepSeek-V4-Pro | DeepSeek | 14% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 24 2026 |
| 69 | DeepSeek-V3.2 | DeepSeek | 13% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Dec 1 2025 |
| 70 | GPT-4o | OpenAI | 12% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | May 13 2024 |
| 70 | gpt-oss-120b | OpenAI | 12% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Aug 5 2025 |
| 72 | Gemini 3.1 Flash-Lite | Google | 11% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Mar 3 2026 |
| 73 | Claude 3 Haiku | Anthropic | 10% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Mar 4 2024 |
| 73 | Kimi K2 | Moonshot AI | 10% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jul 11 2025 |
| 73 | Gemini 3.0 Flash | Google | 10% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Dec 17 2025 |
| 76 | o4-mini | OpenAI | 8% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Apr 16 2025 |
| 76 | DeepSeek R1-0528 | DeepSeek | 8% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | May 28 2025 |
| 76 | GLM-4.5 | Z.ai | 8% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jul 28 2025 |
| 79 | GPT-4o mini | OpenAI | 2% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Jul 18 2024 |


---

Canonical page: https://aireleasetracker.com/benchmark/bullshitbench-v2
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
