Claude 3.7 SonnetvsKimi K3

Claude 3.7 Sonnet
Kimi K3
Specifications
Parameters
2.8T
Context window
1M
Benchmarks
Nonsense detection
BullshitBench v2
49%
73%Best
Coding
SWE-Bench Verified
62.3%
Agentic coding
DeepSWE 1.1
69%
Agentic coding
DeepSWE 1.0
67.5%
Next.js coding
Next.js Evals
92%
Supabase coding
Supabase Evals · with skills
86.4%
Supabase coding
Supabase Evals · no skills
90.9%
Agentic terminal coding
Terminal-Bench 2.1
88.3%
Multi-step tool use
MCP Atlas
84.2%
Professional tool use
JobBench
52.9%
Personal tool use
Toolathlon-Verified
73.2%
Web browsing
BrowseComp
91.2%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
43.5%
Multidisciplinary reasoning
Humanity's Last Exam · with tools
56%
Science
GPQA Diamond
68%
93.5%Best
Knowledge work
GDPval-AA v2
1668
Chart reasoning
CharXiv Reasoning
84.8%
Multimodal reasoning
MMMU-Pro
81.6%
Community preference
Arena Elo (Text)
1486
Community preference (code)
Arena Elo (Code)
1679
Overview
CompanyAnthropicMoonshot AI
Release dateFeb 24 2025Jul 16 2026
AccessProprietaryOpen Weight

Which is better: Claude 3.7 Sonnet or Kimi K3?

Kimi K3 leads Claude 3.7 Sonnet on 2 of the 2 benchmarks they both report (BullshitBench v2, GPQA Diamond). Claude 3.7 Sonnet shipped 507 days before Kimi K3, so benchmark comparisons should account for the intervening progress.

Claude 3.7 Sonnet is proprietary, while Kimi K3 is open weight.

On BullshitBench v2, Kimi K3 leads at 73% vs Claude 3.7 Sonnet at 49%. On GPQA Diamond, Kimi K3 leads at 93.5% vs Claude 3.7 Sonnet at 68%.

Frequently asked questions

Claude 3.7 Sonnet was released by Anthropic on Feb 24 2025.

Kimi K3 was released by Moonshot AI on Jul 16 2026.

Kimi K3 leads on GPQA Diamond — Claude 3.7 Sonnet 68% vs Kimi K3 93.5%.

Claude 3.7 Sonnet is a proprietary model released by Anthropic. Kimi K3 is an open weight model released by Moonshot AI.

Other comparisons