Compare AI models

Specifications
Parameters
—2.8T
Context window
1M1M
API pricing
Input price
$0.75$3.00
Output price
$3.75$15.00
Cached input price
$0.075$0.30
Cheapest input
$0.375Google$1.00InferenceNet
Cheapest output
$1.875Google$7.70Morph
Benchmarks
BullshitBench v2
35%74%
DeepSWE 1.1
65.3%69%
Terminal-Bench 2.1
85.8%88.3%
ARC-AGI-2
84.6%60.4%
GDPval-AA v2
15251668
CharXiv Reasoning
84.5%84.8%
threejseval
14801536
Overview
CompanyGoogleMoonshot AI
Release dateAug 13 2026Jul 16 2026
AccessClosedOpen Weight
Model detailsView modelView model

Frequently asked questions

Kimi K3 leads Gemini 3.7 Flash on 6 of the 7 benchmarks they both report. Gemini 3.7 Flash is cheaper on both input and output: $0.75 vs $3.00 per million input tokens, and $3.75 vs $15.00 per million output tokens. Kimi K3 shipped 28 days before Gemini 3.7 Flash, so benchmark comparisons should account for the intervening progress.

Context windows are 1M (Gemini 3.7 Flash) vs 1M (Kimi K3). Gemini 3.7 Flash is closed, while Kimi K3 is open weight.

On BullshitBench v2, Kimi K3 leads at 74% vs Gemini 3.7 Flash at 35%. On DeepSWE 1.1, Kimi K3 leads at 69% vs Gemini 3.7 Flash at 65.3%. On Terminal-Bench 2.1, Kimi K3 leads at 88.3% vs Gemini 3.7 Flash at 85.8%. On ARC-AGI-2, Gemini 3.7 Flash leads at 84.6% vs Kimi K3 at 60.4%. On GDPval-AA v2, Kimi K3 leads at 1668 vs Gemini 3.7 Flash at 1525. On CharXiv Reasoning, Kimi K3 leads at 84.8% vs Gemini 3.7 Flash at 84.5%. On threejseval, Kimi K3 leads at 1536 vs Gemini 3.7 Flash at 1480.