Claude Opus 4.8vsComposer 2.5

Claude Opus 4.8
Composer 2.5
Specifications
Context window
1M
API pricing
Input price
$5.00
Output price
$25.00
Cached input price
$0.50
Cheapest input
$5.00Amazon Bedrock
Cheapest output
$25.00Amazon Bedrock
Benchmarks
SWE-Bench Pro
69.2%54%
SWE-Bench Multilingual
84.4%79.8%
DeepSWE 1.0
55.8%18%
Next.js Evals
81%81%
Terminal-Bench 2.1
74.6%73%
Benchmarks
BullshitBench v2
95%
Gray Swan IPI · k = 1
0.5%
Gray Swan IPI · k = 10
4.1%
Gray Swan IPI · k = 15
5.5%
SWE-Bench Verified
88.6%
Frontier-Bench v0.1
21.1%
Terminal-Bench 4.0
23.64%
Terminal-Bench 2.0
69.3%
BU Bench
74%
BrowseComp
84.3%
Humanity's Last Exam · no tools
49.8%
Humanity's Last Exam · with tools
57.9%
ARC-AGI-2
72.08%
OSWorld-Verified
83.4%
Finance Agent v2
53.9%
Harvey's Legal Agent Benchmark
9.58%
TaxEval v2
75.63%
MedScribe
85.75%
GDPval-AA
1890
GDPval-AA v2
1600
Overview
CompanyAnthropicSpaceXAI
Release dateMay 28 2026May 18 2026
AccessProprietaryProprietary

Other comparisons

Claude Opus 4.8vsGPT-6 AstraComposer 2.5vsGPT-6 AstraClaude Opus 4.8vsGemini 3.8 FlashComposer 2.5vsGemini 3.8 FlashClaude Opus 4.8vsMuse Spark 1.3Composer 2.5vsMuse Spark 1.3Claude Opus 4.8vsDeepSeek-V4-Pro-0813Composer 2.5vsDeepSeek-V4-Pro-0813Claude Opus 4.8vsMistral Medium 3.5Composer 2.5vsMistral Medium 3.5Claude Opus 4.8vsKimi K3Composer 2.5vsKimi K3

Frequently asked questions

Claude Opus 4.8 leads Composer 2.5 on 4 of the 5 benchmarks they both report. Only Claude Opus 4.8 has a verified first-party API price: $5.00 per million input tokens and $25.00 per million output tokens. No pay-as-you-go API rate is tracked for Composer 2.5. Composer 2.5 shipped 10 days before Claude Opus 4.8, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On SWE-Bench Pro, Claude Opus 4.8 leads at 69.2% vs Composer 2.5 at 54%. On SWE-Bench Multilingual, Claude Opus 4.8 leads at 84.4% vs Composer 2.5 at 79.8%. On DeepSWE 1.0, Claude Opus 4.8 leads at 55.8% vs Composer 2.5 at 18%. On Next.js Evals, both models score 81%. On Terminal-Bench 2.1, Claude Opus 4.8 leads at 74.6% vs Composer 2.5 at 73%.