Kimi K3vsGrok 4.5

Kimi K3
Grok 4.5
Specifications
Parameters
2.8T
Context window
1M
API pricing
Input price
$3.00$2.00
Output price
$15.00$6.00
Cached input price
$0.30$0.30
Cheapest input
$2.10InferenceNet
Cheapest output
$10.95InferenceNet
Benchmarks
BullshitBench v2
74%55%
DeepSWE 1.1
69%54%
DeepSWE 1.0
67.5%62%
Next.js Evals
84%65%
Terminal-Bench 2.1
88.3%83.3%
GDPval-AA v2
16681526
Benchmarks
Gray Swan IPI · k = 1
13.4%
Gray Swan IPI · k = 10
54.2%
Gray Swan IPI · k = 15
60.8%
SWE-Bench Pro
64.7%
SWE-Bench Multilingual
78%
FrontierCode v1.1 (Extended) · extended split
56.6%
APEX-SWE
53.6%
Supabase Evals · with skills
84.1%
Supabase Evals · no skills
78.3%
Frontier-Bench v0.1
17.8%
Terminal-Bench 4.0
12.42%
Terminal-Bench 3.0
15.7%
APEX-Agents
47.1%
MCP Atlas
84.2%
JobBench
52.9%
Toolathlon-Verified
73.2%
BrowseComp
91.2%
Humanity's Last Exam · no tools
43.5%
Humanity's Last Exam · with tools
56%
ARC-AGI-2
52.64%
GPQA Diamond
93.5%
Harvey's Legal Agent Benchmark
12.92%
MedScribe
86.88%
AA Intelligence Index
56
AA-Briefcase
1313
CharXiv Reasoning
84.8%
MMMU-Pro
81.6%
threejseval
1558
Overview
CompanyMoonshot AISpaceXAI
Release dateJul 16 2026Jul 8 2026
AccessOpen WeightProprietary

Other comparisons

Kimi K3vsClaude Fable 5.1Grok 4.5vsClaude Fable 5.1Kimi K3vsGPT-6 AstraGrok 4.5vsGPT-6 AstraKimi K3vsGemini 3.8 FlashGrok 4.5vsGemini 3.8 FlashKimi K3vsMuse Spark 1.3Grok 4.5vsMuse Spark 1.3Kimi K3vsDeepSeek-V4.1-FlashGrok 4.5vsDeepSeek-V4.1-FlashKimi K3vsMistral Medium 3.5Grok 4.5vsMistral Medium 3.5

Frequently asked questions

Kimi K3 leads Grok 4.5 on 6 of the 6 benchmarks they both report. Grok 4.5 is cheaper on both input and output: $2.00 vs $3.00 per million input tokens, and $6.00 vs $15.00 per million output tokens. Figures are base-tier rates. Grok 4.5 shipped 8 days before Kimi K3, so benchmark comparisons should account for the intervening progress.

Kimi K3 is open weight, while Grok 4.5 is proprietary.

On BullshitBench v2, Kimi K3 leads at 74% vs Grok 4.5 at 55%. On DeepSWE 1.1, Kimi K3 leads at 69% vs Grok 4.5 at 54%. On DeepSWE 1.0, Kimi K3 leads at 67.5% vs Grok 4.5 at 62%. On Next.js Evals, Kimi K3 leads at 84% vs Grok 4.5 at 65%. On Terminal-Bench 2.1, Kimi K3 leads at 88.3% vs Grok 4.5 at 83.3%. On GDPval-AA v2, Kimi K3 leads at 1668 vs Grok 4.5 at 1526.