Kimi K2.5vsQwen3.5

Kimi K2.5
Qwen3.5
Specifications
Parameters
1T
397B
Context window
256k
1M
Benchmarks
Nonsense detection
BullshitBench v2
52%
Coding
SWE-Bench Verified
76.8%Best
76.4%
Multilingual coding
SWE-Bench Multilingual
69.3%
Agentic coding
CursorBench v3.1
31.9%
Next.js coding
Next.js Evals
21%
Competitive coding
LiveCodeBench
85%
Agentic terminal coding
Terminal-Bench 2.0
52.5%
Web browsing
BrowseComp
69%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
28.7%
Science
GPQA Diamond
87.6%
88.4%Best
Agentic computer use
OSWorld-Verified
62.2%
Chart reasoning
CharXiv Reasoning
80.8%
Multimodal reasoning
MMMU-Pro
79%
Multimodal
MMMU
85%
Community preference (code)
Arena Elo (Code)
1433
Overview
CompanyMoonshot AIQwen
Release dateJan 27 2026Feb 16 2026
AccessOpen WeightOpen Weight

Which is better: Kimi K2.5 or Qwen3.5?

Kimi K2.5 and Qwen3.5 are evenly matched across the 2 benchmarks they both report (SWE-Bench Verified, GPQA Diamond). Kimi K2.5 shipped 20 days before Qwen3.5, so benchmark comparisons should account for the intervening progress.

Kimi K2.5 has 1T parameters, while Qwen3.5 has 397B. Context windows are 256k (Kimi K2.5) vs 1M (Qwen3.5).

On SWE-Bench Verified, Kimi K2.5 leads at 76.8% vs Qwen3.5 at 76.4%. On GPQA Diamond, Qwen3.5 leads at 88.4% vs Kimi K2.5 at 87.6%.

Frequently asked questions

Kimi K2.5 was released by Moonshot AI on Jan 27 2026.

Qwen3.5 was released by Qwen on Feb 16 2026.

Kimi K2.5 leads on SWE-Bench Verified — Kimi K2.5 76.8% vs Qwen3.5 76.4%.

Qwen3.5 leads on GPQA Diamond — Kimi K2.5 87.6% vs Qwen3.5 88.4%.

Kimi K2.5 has a 256k context window; Qwen3.5 has 1M.

Other comparisons