Claude Opus 4.8vsKimi K3

Claude Opus 4.8
Kimi K3
Specifications
Parameters
2.8T
Context window
1M1M
API pricing
Input price
$5.00$3.00
Output price
$25.00$15.00
Cached input price
$0.50$0.30
Cheapest input
$5.00Amazon Bedrock$2.125Morph
Cheapest output
$25.00Amazon Bedrock$11.5502Sail Research
Benchmarks
BullshitBench v2
95%74%
DeepSWE 1.0
55.8%67.5%
Next.js Evals
68%84%
Terminal-Bench 2.1
74.6%88.3%
BrowseComp
84.3%91.2%
Humanity's Last Exam · no tools
49.8%43.5%
Humanity's Last Exam · with tools
57.9%56%
GDPval-AA v2
16001668
Benchmarks
Gray Swan IPI · k = 1
0.5%
Gray Swan IPI · k = 10
4.1%
Gray Swan IPI · k = 15
5.5%
ProgramBench
0%
SWE-Bench Pro
69.2%
SWE-Bench Verified
88.6%
SWE-Bench Multilingual
84.4%
DeepSWE 1.1
69%
Supabase Evals · with skills
78.3%
Supabase Evals · no skills
81.2%
Frontier-Bench v0.1
21.1%
Terminal-Bench 4.0
23.64%
MCP Atlas
84.2%
JobBench
52.9%
Toolathlon-Verified
73.2%
BU Bench
74%
ARC-AGI-2
72.08%
GPQA Diamond
93.5%
OSWorld-Verified
83.4%
Finance Agent v2
53.9%
Harvey's Legal Agent Benchmark
9.58%
TaxEval v2
75.63%
MedScribe
85.75%
GDPval-AA
1890
CharXiv Reasoning
84.8%
MMMU-Pro
81.6%
threejseval
1549
Overview
CompanyAnthropicMoonshot AI
Release dateMay 28 2026Jul 16 2026
AccessProprietaryOpen Weight

Other comparisons

Claude Opus 4.8vsGPT-6 AstraKimi K3vsGPT-6 AstraClaude Opus 4.8vsGemini 3.8 FlashKimi K3vsGemini 3.8 FlashClaude Opus 4.8vsMuse Spark 1.3Kimi K3vsMuse Spark 1.3Claude Opus 4.8vsGrok 4.6Kimi K3vsGrok 4.6Claude Opus 4.8vsDeepSeek-V4.1-FlashKimi K3vsDeepSeek-V4.1-FlashClaude Opus 4.8vsMistral Medium 3.5Kimi K3vsMistral Medium 3.5

Frequently asked questions

Kimi K3 leads Claude Opus 4.8 on 5 of the 8 benchmarks they both report. Kimi K3 is cheaper on both input and output: $3.00 vs $5.00 per million input tokens, and $15.00 vs $25.00 per million output tokens. Claude Opus 4.8 shipped 49 days before Kimi K3, so benchmark comparisons should account for the intervening progress.

Context windows are 1M (Claude Opus 4.8) vs 1M (Kimi K3). Claude Opus 4.8 is proprietary, while Kimi K3 is open weight.

On BullshitBench v2, Claude Opus 4.8 leads at 95% vs Kimi K3 at 74%. On DeepSWE 1.0, Kimi K3 leads at 67.5% vs Claude Opus 4.8 at 55.8%. On Next.js Evals, Kimi K3 leads at 84% vs Claude Opus 4.8 at 68%. On Terminal-Bench 2.1, Kimi K3 leads at 88.3% vs Claude Opus 4.8 at 74.6%. On BrowseComp, Kimi K3 leads at 91.2% vs Claude Opus 4.8 at 84.3%. On Humanity's Last Exam · no tools, Claude Opus 4.8 leads at 49.8% vs Kimi K3 at 43.5%. On Humanity's Last Exam · with tools, Claude Opus 4.8 leads at 57.9% vs Kimi K3 at 56%. On GDPval-AA v2, Kimi K3 leads at 1668 vs Claude Opus 4.8 at 1600.