Gemini 3.1 ProvsKimi K3

Gemini 3.1 Pro
Kimi K3
Specifications
Parameters
2.8T
Context window
1M
API pricing
Input price
$2.00$3.00
Output price
$12.00$15.00
Cached input price
$0.20$0.30
Cheapest input
$2.10InferenceNet
Cheapest output
$10.95InferenceNet
Benchmarks
BullshitBench v2
37%74%
DeepSWE 1.1
12%69%
Next.js Evals
58%84%
Terminal-Bench 2.1
70.3%88.3%
MCP Atlas
78.2%84.2%
BrowseComp
85.9%91.2%
Humanity's Last Exam · no tools
44.4%43.5%
Humanity's Last Exam · with tools
51.4%56%
GPQA Diamond
94.3%93.5%
GDPval-AA v2
9651668
CharXiv Reasoning
83.3%84.8%
MMMU-Pro
80.5%81.6%
Benchmarks
Gray Swan IPI · k = 1
14.2%
Gray Swan IPI · k = 10
45.7%
Gray Swan IPI · k = 15
49.2%
ProgramBench
0%
SWE-Bench Pro
54.2%
SWE-Bench Verified
80.6%
DeepSWE 1.0
67.5%
MLE-Bench
42.6%
Supabase Evals · with skills
84.1%
Supabase Evals · no skills
78.3%
Terminal-Bench 2.0
68.5%
JobBench
52.9%
Toolathlon-Verified
73.2%
Toolathlon
48.8%
ARC-AGI-2
77.1%
FrontierMath · Tier 1–3
36.9%
FrontierMath · Tier 4
16.7%
OSWorld-Verified
76.2%
Finance Agent v2
43%
GDPval-AA
1314
GDPval (win/tie rate)
67.3%
Blueprint-Bench 2
26.5%
MRCR v2 (8-needle) · 128k average
84.9%
MRCR v2 (8-needle) · 1M pointwise
26.3%
threejseval
1558
Overview
CompanyGoogleMoonshot AI
Release dateFeb 19 2026Jul 16 2026
AccessProprietaryOpen Weight

Other comparisons

Gemini 3.1 ProvsClaude Fable 5.1Kimi K3vsClaude Fable 5.1Gemini 3.1 ProvsGPT-6 AstraKimi K3vsGPT-6 AstraGemini 3.1 ProvsMuse Spark 1.3Kimi K3vsMuse Spark 1.3Gemini 3.1 ProvsGrok 4.6Kimi K3vsGrok 4.6Gemini 3.1 ProvsDeepSeek-V4.1-FlashKimi K3vsDeepSeek-V4.1-FlashGemini 3.1 ProvsMistral Medium 3.5Kimi K3vsMistral Medium 3.5

Frequently asked questions

Kimi K3 leads Gemini 3.1 Pro on 10 of the 12 benchmarks they both report. Gemini 3.1 Pro is cheaper on both input and output: $2.00 vs $3.00 per million input tokens, and $12.00 vs $15.00 per million output tokens. Figures are base-tier rates. Gemini 3.1 Pro shipped 147 days before Kimi K3, so benchmark comparisons should account for the intervening progress.

Gemini 3.1 Pro is proprietary, while Kimi K3 is open weight.

On BullshitBench v2, Kimi K3 leads at 74% vs Gemini 3.1 Pro at 37%. On DeepSWE 1.1, Kimi K3 leads at 69% vs Gemini 3.1 Pro at 12%. On Next.js Evals, Kimi K3 leads at 84% vs Gemini 3.1 Pro at 58%. On Terminal-Bench 2.1, Kimi K3 leads at 88.3% vs Gemini 3.1 Pro at 70.3%. On MCP Atlas, Kimi K3 leads at 84.2% vs Gemini 3.1 Pro at 78.2%. On BrowseComp, Kimi K3 leads at 91.2% vs Gemini 3.1 Pro at 85.9%. On Humanity's Last Exam · no tools, Gemini 3.1 Pro leads at 44.4% vs Kimi K3 at 43.5%. On Humanity's Last Exam · with tools, Kimi K3 leads at 56% vs Gemini 3.1 Pro at 51.4%. On GPQA Diamond, Gemini 3.1 Pro leads at 94.3% vs Kimi K3 at 93.5%. On GDPval-AA v2, Kimi K3 leads at 1668 vs Gemini 3.1 Pro at 965. On CharXiv Reasoning, Kimi K3 leads at 84.8% vs Gemini 3.1 Pro at 83.3%. On MMMU-Pro, Kimi K3 leads at 81.6% vs Gemini 3.1 Pro at 80.5%.