Compare AI models

Specifications
Parameters
—125B
Context window
—262k
Benchmarks
SWE-Bench Pro
55%62.5%
DeepSWE 1.1
10%58.7%
JobBench
17%55.7%
Toolathlon-Verified
49.4%73.5%
Humanity's Last Exam · with tools
50.4%35.9%
GPQA Diamond
89.5%91.7%
CharXiv Reasoning
88.9%84.6%
Overview
CompanyMetaQwen
Release dateApr 8 2026Aug 26 2026
AccessClosedOpen Weight
Model detailsView modelView model

Frequently asked questions

Qwen3.8-Flash-Next leads Muse Spark on 5 of the 7 benchmarks they both report. Muse Spark shipped 140 days before Qwen3.8-Flash-Next, so benchmark comparisons should account for the intervening progress.

Muse Spark is closed, while Qwen3.8-Flash-Next is open weight.

On SWE-Bench Pro, Qwen3.8-Flash-Next leads at 62.5% vs Muse Spark at 55%. On DeepSWE 1.1, Qwen3.8-Flash-Next leads at 58.7% vs Muse Spark at 10%. On JobBench, Qwen3.8-Flash-Next leads at 55.7% vs Muse Spark at 17%. On Toolathlon-Verified, Qwen3.8-Flash-Next leads at 73.5% vs Muse Spark at 49.4%. On Humanity's Last Exam · with tools, Muse Spark leads at 50.4% vs Qwen3.8-Flash-Next at 35.9%. On GPQA Diamond, Qwen3.8-Flash-Next leads at 91.7% vs Muse Spark at 89.5%. On CharXiv Reasoning, Muse Spark leads at 88.9% vs Qwen3.8-Flash-Next at 84.6%.