Anthropic
Claude 3.5 Haiku
Claude 3.5 Haiku is an AI model released by Anthropic on Tuesday, Oct 22 2024, 124 days after Claude 3.5 Sonnet. Benchmark results (shown below) cover BullshitBench v2, SWE-Bench Verified, and GPQA Diamond.
Benchmarks
Nonsense detection
BullshitBench v2Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
50%
#16 of 60Best: Claude Opus 4.8 · 95%
Coding
SWE-Bench VerifiedReal coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better.
40.6%
#31 of 33Best: Claude Fable 5 · 95.5%
Science
GPQA DiamondGraduate-level science questions in biology, physics, and chemistry — hard enough that subject-matter PhDs score around 65%. Higher is better.
41.6%
#35 of 37Best: GPT-5.4-Pro · 94.4%
Compare Claude 3.5 Haiku with
Claude 3.5 Haiku vs Claude Sonnet 5Claude 3.5 Haiku vs GPT-5.6 SolClaude 3.5 Haiku vs Gemini 3.5 FlashClaude 3.5 Haiku vs Muse Spark 1.1Claude 3.5 Haiku vs Grok 4.5Claude 3.5 Haiku vs DeepSeek-V4-ProClaude 3.5 Haiku vs Mistral Medium 3.5Claude 3.5 Haiku vs Kimi K3Claude 3.5 Haiku vs Composer 2.5Claude 3.5 Haiku vs GLM-5.2