Claude Opus 5vso3

Claude Opus 5
o3
Specifications
Context window
1M
API pricing
Input price
$5.00$2.00
Output price
$25.00$8.00
Cached input price
$0.50$0.50
Cheapest input
$5.00Amazon Bedrock
Cheapest output
$25.00Amazon Bedrock
Benchmarks
BullshitBench v2
73%26%
Benchmarks
Gray Swan IPI · k = 1
0.2%
Gray Swan IPI · k = 10
1.6%
Gray Swan IPI · k = 15
2%
SWE-Bench Pro
79.2%
SWE-Bench Verified
96%
SWE-Bench Multilingual
89.5%
SWE-Bench Multimodal
59.4%
DeepSWE 1.1
68.8%
FrontierCode v1.1 (Main) · main split
53.4%
Next.js Evals
92%
Supabase Evals · with skills
91.3%
Supabase Evals · no skills
89.9%
Frontier-Bench v0.1
43.3%
Terminal-Bench 4.0
51.82%
Terminal-Bench-Science 0.1
29%
BrowseComp
90.8%
Humanity's Last Exam · no tools
56.3%
Humanity's Last Exam · with tools
64.7%
ARC-AGI-3
30.2%
ARC-AGI-2
90.4%
BioMysteryBench · hard
49.4%
BioMysteryBench · human solved
90.1%
OSWorld 2.0
70.6%
AutomationBench
26%
Harvey's Legal Agent Benchmark (Held-out)
11.7%
HealthBench Professional
59.8%
GDPval-AA v2
1861
AA-Briefcase
1685
threejseval
1770
Overview
CompanyAnthropicOpenAI
Release dateJul 24 2026Apr 16 2025
AccessProprietaryProprietary

Other comparisons

Claude Opus 5vsGemini 3.8 Flasho3vsGemini 3.8 FlashClaude Opus 5vsMuse Spark 1.3o3vsMuse Spark 1.3Claude Opus 5vsGrok 4.6o3vsGrok 4.6Claude Opus 5vsDeepSeek-V4-Pro-0813o3vsDeepSeek-V4-Pro-0813Claude Opus 5vsMistral Medium 3.5o3vsMistral Medium 3.5Claude Opus 5vsKimi K3o3vsKimi K3

Frequently asked questions

Claude Opus 5 leads o3 on 1 of the 1 benchmark they both report (BullshitBench v2). o3 is cheaper on both input and output: $2.00 vs $5.00 per million input tokens, and $8.00 vs $25.00 per million output tokens. o3 shipped 464 days before Claude Opus 5, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude Opus 5 leads at 73% vs o3 at 26%.