Claude Opus 5vsKimi K2 Thinking

Claude Opus 5
Kimi K2 Thinking
Specifications
Parameters
1T
Context window
1M256k
API pricing
Input price
$5.00
Output price
$25.00
Cached input price
$0.50
Cheapest input
$5.00Amazon Bedrock$0.60Google
Cheapest output
$25.00Amazon Bedrock$2.50Google
Benchmarks
SWE-Bench Verified
96%71.3%
Benchmarks
BullshitBench v2
73%
Gray Swan IPI · k = 1
0.2%
Gray Swan IPI · k = 10
1.6%
Gray Swan IPI · k = 15
2%
SWE-Bench Pro
79.2%
SWE-Bench Multilingual
89.5%
SWE-Bench Multimodal
59.4%
DeepSWE 1.1
68.8%
FrontierCode v1.1 (Main) · main split
53.4%
Next.js Evals
92%
Supabase Evals · with skills
91.3%
Supabase Evals · no skills
89.9%
Frontier-Bench v0.1
43.3%
Terminal-Bench 4.0
51.82%
Terminal-Bench-Science 0.1
29%
BrowseComp
90.8%
Humanity's Last Exam · no tools
56.3%
Humanity's Last Exam · with tools
64.7%
ARC-AGI-3
30.2%
ARC-AGI-2
90.4%
BioMysteryBench · hard
49.4%
BioMysteryBench · human solved
90.1%
OSWorld 2.0
70.6%
AutomationBench
26%
Harvey's Legal Agent Benchmark (Held-out)
11.7%
HealthBench Professional
59.8%
GDPval-AA v2
1861
AA-Briefcase
1685
threejseval
1770
Overview
CompanyAnthropicMoonshot AI
Release dateJul 24 2026Nov 6 2025
AccessProprietaryOpen Weight

Other comparisons

Claude Opus 5vsGPT-6 AstraKimi K2 ThinkingvsGPT-6 AstraClaude Opus 5vsGemini 3.8 FlashKimi K2 ThinkingvsGemini 3.8 FlashClaude Opus 5vsMuse Spark 1.3Kimi K2 ThinkingvsMuse Spark 1.3Claude Opus 5vsGrok 4.6Kimi K2 ThinkingvsGrok 4.6Claude Opus 5vsDeepSeek-V4-Pro-0813Kimi K2 ThinkingvsDeepSeek-V4-Pro-0813Claude Opus 5vsMistral Medium 3.5Kimi K2 ThinkingvsMistral Medium 3.5

Frequently asked questions

Claude Opus 5 leads Kimi K2 Thinking on 1 of the 1 benchmark they both report (SWE-Bench Verified). Only Claude Opus 5 has a verified first-party API price: $5.00 per million input tokens and $25.00 per million output tokens. No pay-as-you-go API rate is tracked for Kimi K2 Thinking. Kimi K2 Thinking shipped 260 days before Claude Opus 5, so benchmark comparisons should account for the intervening progress.

Context windows are 1M (Claude Opus 5) vs 256k (Kimi K2 Thinking). Claude Opus 5 is proprietary, while Kimi K2 Thinking is open weight.

On SWE-Bench Verified, Claude Opus 5 leads at 96% vs Kimi K2 Thinking at 71.3%.