Claude Opus 5vsKimi K2.5
Claude Opus 5 | Kimi K2.5 | |
|---|---|---|
| Specifications | ||
Parameters | — | 1T |
Context window | 1M | 256k |
| Benchmarks | ||
Nonsense detection BullshitBench v2Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better. | — | 52% |
Coding SWE-Bench VerifiedReal coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better. | — | 76.8% |
Agentic coding CursorBench v3.1Cursor's own test of harder, real-world coding tasks inside a code editor. Higher is better. | — | 31.9% |
Agentic coding DeepSWE 1.1Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. | 68.8% | — |
Agentic coding FrontierCode v1.1 (Main)A set of very hard, frontier-difficulty coding tasks an AI agent has to complete end to end. The score is the share of tasks in the main split it solves. Higher is better. | 53.4% | — |
Next.js coding Next.js EvalsVercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better. | — | 21% |
Competitive coding LiveCodeBenchCoding problems published so recently the AI can't have seen them in training — a contamination-free test of raw programming skill. Higher is better. | — | 85% |
Agentic computer work Frontier-Bench v0.1A hard, ever-evolving set of real computer tasks — coding, system administration, data work, and more — that an AI agent has to complete on its own. Run by the Harbor / Laude Institute team as the successor to Terminal-Bench (v0.1 is the first release of the task set). The score is the share of tasks solved. Higher is better. | 43.3% | — |
Web browsing BrowseCompCan the AI browse the web and track down hard-to-find answers? Higher is better. | 90.8% | — |
Multidisciplinary reasoning Humanity's Last Exam · no toolsHumanity's Last Exam — extremely hard expert questions across many subjects, written so you can't just look up the answer. “No tools” means the AI answers on its own. Higher is better. | 56.3% | — |
Multidisciplinary reasoning Humanity's Last Exam · with toolsHumanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better. | 64.7% | — |
Novel problem-solving ARC-AGI-3The third generation of the ARC-AGI series: instead of static puzzles, the AI is dropped into small interactive game-like environments it has never seen and has to figure out the rules and solve them on its own. Higher is better. | 30.2% | — |
Biology BioMysteryBench · hardReal unsolved-style biology puzzles — the AI has to reason its way to an answer the way a research biologist would. The “hard” split contains the toughest cases. Higher is better. | 49.4% | — |
Biology BioMysteryBench · human solvedReal biology puzzles that human experts have managed to crack — can the AI reach the same answers? Higher is better. | 90.1% | — |
Science GPQA DiamondGraduate-level science questions in biology, physics, and chemistry — hard enough that subject-matter PhDs score around 65%. Higher is better. | — | 87.6% |
Agentic computer use OSWorld 2.0Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Version 2.0 is a harder, refreshed task set. Higher is better. | 70.6% | — |
Business workflows AutomationBenchTests whether the AI can run real multi-step business workflows — the kind of end-to-end office processes companies want to automate — from start to finish. Higher is better. | 26% | — |
Agentic legal work Harvey's Legal Agent Benchmark (Held-out)Harvey's test of whether an AI agent can complete real legal work, scored on a held-out set of tasks the model makers never see — making the numbers harder to game. Higher is better. | 11.7% | — |
Health HealthBench ProfessionalRealistic health conversations graded against detailed rubrics written by physicians — can the AI respond the way a careful medical professional would? Higher is better. | 59.8% | — |
Knowledge work GDPval-AA v2economically valuable knowledge work (v2, re-based Elo) | 1861 | — |
Community preference (code) Arena Elo (Code)Like the text arena, but people vote on which AI writes better code. The votes become a chess-style Elo rating on arena.ai. Higher is better. | — | 1433 |
| Overview | ||
| Company | Anthropic | Moonshot AI |
| Release date | Jul 24 2026 | Jan 27 2026 |
| Access | Proprietary | Open Weight |
Which is better: Claude Opus 5 or Kimi K2.5?
Claude Opus 5 and Kimi K2.5 don't publish scores on any of the same benchmarks, so there's no direct head-to-head comparison. Kimi K2.5 shipped 178 days before Claude Opus 5, so benchmark comparisons should account for the intervening progress.
Context windows are 1M (Claude Opus 5) vs 256k (Kimi K2.5). Claude Opus 5 is proprietary, while Kimi K2.5 is open weight.
Direct benchmark comparisons are unavailable — Claude Opus 5 and Kimi K2.5 don't publish scores on any of the same benchmarks.
Frequently asked questions
Claude Opus 5 was released by Anthropic on Jul 24 2026.
Kimi K2.5 was released by Moonshot AI on Jan 27 2026.
Claude Opus 5 has a 1M context window; Kimi K2.5 has 256k.
Claude Opus 5 is a proprietary model released by Anthropic. Kimi K2.5 is an open weight model released by Moonshot AI.