Claude 3.5 SonnetvsGPT-4o

Claude 3.5 Sonnet
GPT-4o
Specifications
Context window
128k
Benchmarks
Nonsense detection
BullshitBench v2
45%Best
12%
Coding
SWE-Bench Verified
33.4%
Science
GPQA Diamond
59.4%Best
49.9%
Overview
CompanyAnthropicOpenAI
Release dateJun 20 2024May 13 2024
AccessProprietaryProprietary

Which is better: Claude 3.5 Sonnet or GPT-4o?

Claude 3.5 Sonnet leads GPT-4o on 2 of the 2 benchmarks they both report (BullshitBench v2, GPQA Diamond). GPT-4o shipped 38 days before Claude 3.5 Sonnet, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude 3.5 Sonnet leads at 45% vs GPT-4o at 12%. On GPQA Diamond, Claude 3.5 Sonnet leads at 59.4% vs GPT-4o at 49.9%.

Frequently asked questions

Claude 3.5 Sonnet was released by Anthropic on Jun 20 2024.

GPT-4o was released by OpenAI on May 13 2024.

Claude 3.5 Sonnet leads on GPQA Diamond — Claude 3.5 Sonnet 59.4% vs GPT-4o 49.9%.

Other comparisons