Claude Opus 4.6vsQwen3.5

Claude Opus 4.6
Qwen3.5
Specifications
Parameters
397B
Context window
1M
Benchmarks
Nonsense detection
BullshitBench v2
87%
Coding
SWE-Bench Verified
80.8%Best
76.4%
Multilingual coding
SWE-Bench Multilingual
69.3%
Next.js coding
Next.js Evals
75%
Agentic terminal coding
Terminal-Bench 2.0
52.5%
Web browsing
BrowseComp
83.7%Best
69%
Multidisciplinary reasoning
Humanity's Last Exam · no tools
28.7%
Multidisciplinary reasoning
Humanity's Last Exam · with tools
53%
Science
GPQA Diamond
91.3%Best
88.4%
Agentic computer use
OSWorld-Verified
62.2%
Chart reasoning
CharXiv Reasoning
80.8%
Multimodal reasoning
MMMU-Pro
79%
Multimodal
MMMU
85%
Community preference
Arena Elo (Text)
1504
Community preference (code)
Arena Elo (Code)
1543
Overview
CompanyAnthropicQwen
Release dateFeb 5 2026Feb 16 2026
AccessProprietaryOpen Weight

Which is better: Claude Opus 4.6 or Qwen3.5?

Claude Opus 4.6 leads Qwen3.5 on 3 of the 3 benchmarks they both report (SWE-Bench Verified, BrowseComp, GPQA Diamond). Claude Opus 4.6 shipped 11 days before Qwen3.5, so benchmark comparisons should account for the intervening progress.

Claude Opus 4.6 is proprietary, while Qwen3.5 is open weight.

On SWE-Bench Verified, Claude Opus 4.6 leads at 80.8% vs Qwen3.5 at 76.4%. On BrowseComp, Claude Opus 4.6 leads at 83.7% vs Qwen3.5 at 69%. On GPQA Diamond, Claude Opus 4.6 leads at 91.3% vs Qwen3.5 at 88.4%.

Frequently asked questions

Claude Opus 4.6 was released by Anthropic on Feb 5 2026.

Qwen3.5 was released by Qwen on Feb 16 2026.

Claude Opus 4.6 leads on SWE-Bench Verified — Claude Opus 4.6 80.8% vs Qwen3.5 76.4%.

Claude Opus 4.6 leads on GPQA Diamond — Claude Opus 4.6 91.3% vs Qwen3.5 88.4%.

Claude Opus 4.6 is a proprietary model released by Anthropic. Qwen3.5 is an open weight model released by Qwen.

Other comparisons