gpt-oss-20bvsQwen3.8-Flash-Next

gpt-oss-20b
Qwen3.8-Flash-Next
Specifications
Parameters
21B125B
Context window
128k262k
API pricing
Cheapest input
$0.02Darkbloom
Cheapest output
$0.10Darkbloom
Benchmarks
Humanity's Last Exam · no tools
10.9%35.9%
GPQA Diamond
71.5%91.7%
Benchmarks
SWE-Bench Pro
62.5%
SWE-Bench Verified
60.7%
SWE-Bench Multilingual
81%
DeepSWE 1.1
58.7%
NL2Repo-Bench
48.1%
LiveCodeBench
91.9%
JobBench
55.7%
CoWorkBench
73.9%
Toolathlon-Verified
73.5%
Humanity's Last Exam · with tools
17.3%
IFBench
81.3%
MMLU
85.3%
OSWorld 2.0
19.4%
Agent's Last Exam · pass@1
24.3%
Agent's Last Exam · score
51.2%
CharXiv Reasoning
84.6%
LVBench
76.6%
Overview
CompanyOpenAIQwen
Release dateAug 5 2025Aug 26 2026
AccessOpen WeightOpen Weight

Other comparisons

gpt-oss-20bvsClaude Opus 5Qwen3.8-Flash-NextvsClaude Opus 5gpt-oss-20bvsGemini 3.7 FlashQwen3.8-Flash-NextvsGemini 3.7 Flashgpt-oss-20bvsMuse GlimmerQwen3.8-Flash-NextvsMuse Glimmergpt-oss-20bvsGrok 4.6Qwen3.8-Flash-NextvsGrok 4.6gpt-oss-20bvsDeepSeek-V4-Pro-0813Qwen3.8-Flash-NextvsDeepSeek-V4-Pro-0813gpt-oss-20bvsMistral Medium 3.5Qwen3.8-Flash-NextvsMistral Medium 3.5

Frequently asked questions

Qwen3.8-Flash-Next leads gpt-oss-20b on 2 of the 2 benchmarks they both report (Humanity's Last Exam, GPQA Diamond). gpt-oss-20b shipped 386 days before Qwen3.8-Flash-Next, so benchmark comparisons should account for the intervening progress.

gpt-oss-20b has 21B parameters, while Qwen3.8-Flash-Next has 125B. Context windows are 128k (gpt-oss-20b) vs 262k (Qwen3.8-Flash-Next).

On Humanity's Last Exam · no tools, Qwen3.8-Flash-Next leads at 35.9% vs gpt-oss-20b at 10.9%. On GPQA Diamond, Qwen3.8-Flash-Next leads at 91.7% vs gpt-oss-20b at 71.5%.

gpt-oss-20b was released by OpenAI on Aug 5 2025.

Qwen3.8-Flash-Next was released by Qwen on Aug 26 2026.

Qwen3.8-Flash-Next leads on Humanity's Last Exam · no tools — gpt-oss-20b 10.9% vs Qwen3.8-Flash-Next 35.9%.

gpt-oss-20b has a 128k context window; Qwen3.8-Flash-Next has 262k.