Claude Opus 5vsKimi K2.5

Claude Opus 5
Kimi K2.5
Specifications
Parameters
1T
Context window
1M256k
API pricing
Input price
$5.00$0.60
Output price
$25.00$3.00
Cached input price
$0.50$0.10
Cheapest input
$5.00Amazon Bedrock$0.45SiliconFlow
Cheapest output
$25.00Amazon Bedrock$2.25SiliconFlow
Benchmarks
BullshitBench v2
73%52%
SWE-Bench Verified
96%76.8%
Next.js Evals
94%16%
BrowseComp
90.8%60.6%
Humanity's Last Exam · with tools
64.7%30.1%
Benchmarks
Gray Swan IPI · k = 1
0.2%
Gray Swan IPI · k = 10
1.6%
Gray Swan IPI · k = 15
2%
ProgramBench
4.5%
SWE-Bench Pro
79.2%
SWE-Bench Multilingual
89.5%
SWE-Bench Multimodal
59.4%
DeepSWE 1.1
68.8%
FrontierCode v1.1 (Main) · main split
53.4%
Supabase Evals · with skills
91.3%
Supabase Evals · no skills
89.9%
LiveCodeBench
85%
Frontier-Bench v0.1
43.3%
Terminal-Bench 4.0
51.82%
Terminal-Bench-Science 0.1
29%
Humanity's Last Exam · no tools
56.3%
ARC-AGI-3
30.2%
ARC-AGI-2
90.4%
BioMysteryBench · hard
49.4%
BioMysteryBench · human solved
90.1%
GPQA Diamond
87.6%
OSWorld 2.0
70.6%
AutomationBench
26%
Harvey's Legal Agent Benchmark (Held-out)
11.7%
HealthBench Professional
59.8%
GDPval-AA v2
1861
AA-Briefcase
1685
threejseval
1770
Overview
CompanyAnthropicMoonshot AI
Release dateJul 24 2026Jan 27 2026
AccessProprietaryOpen Weight

Other comparisons

Claude Opus 5vsGPT-6 AstraKimi K2.5vsGPT-6 AstraClaude Opus 5vsGemini 3.8 FlashKimi K2.5vsGemini 3.8 FlashClaude Opus 5vsMuse Spark 1.3Kimi K2.5vsMuse Spark 1.3Claude Opus 5vsGrok 4.6Kimi K2.5vsGrok 4.6Claude Opus 5vsDeepSeek-V4.1-FlashKimi K2.5vsDeepSeek-V4.1-FlashClaude Opus 5vsMistral Medium 3.5Kimi K2.5vsMistral Medium 3.5

Frequently asked questions

Claude Opus 5 leads Kimi K2.5 on 5 of the 5 benchmarks they both report. Kimi K2.5 is cheaper on both input and output: $0.60 vs $5.00 per million input tokens, and $3.00 vs $25.00 per million output tokens. Kimi K2.5 shipped 178 days before Claude Opus 5, so benchmark comparisons should account for the intervening progress.

Context windows are 1M (Claude Opus 5) vs 256k (Kimi K2.5). Claude Opus 5 is proprietary, while Kimi K2.5 is open weight.

On BullshitBench v2, Claude Opus 5 leads at 73% vs Kimi K2.5 at 52%. On SWE-Bench Verified, Claude Opus 5 leads at 96% vs Kimi K2.5 at 76.8%. On Next.js Evals, Claude Opus 5 leads at 94% vs Kimi K2.5 at 16%. On BrowseComp, Claude Opus 5 leads at 90.8% vs Kimi K2.5 at 60.6%. On Humanity's Last Exam · with tools, Claude Opus 5 leads at 64.7% vs Kimi K2.5 at 30.1%.