DeepSeek-V4-Provsgpt-oss-120b

DeepSeek-V4-Pro
gpt-oss-120b
Specifications
Parameters
117B
Context window
128k
API pricing
Cheapest input
$0.7078StreamLake$0.03AkashML
Cheapest output
$1.4157StreamLake$0.17AkashML
Benchmarks
BullshitBench v2
14%11%
GPQA Diamond
90.1%80.1%
Benchmarks
SWE-Bench Verified
62.4%
LiveCodeBench
93.5%
BrowseComp
83.4%
Humanity's Last Exam · no tools
14.9%
Humanity's Last Exam · with tools
19%
MMLU
90%
Overview
CompanyDeepSeekOpenAI
Release dateApr 24 2026Aug 5 2025
AccessOpen WeightOpen Weight

Other comparisons

DeepSeek-V4-ProvsClaude Fable 5.1gpt-oss-120bvsClaude Fable 5.1DeepSeek-V4-ProvsGemini 3.8 Flashgpt-oss-120bvsGemini 3.8 FlashDeepSeek-V4-ProvsMuse Spark 1.3gpt-oss-120bvsMuse Spark 1.3DeepSeek-V4-ProvsGrok 4.6gpt-oss-120bvsGrok 4.6DeepSeek-V4-ProvsMistral Medium 3.5gpt-oss-120bvsMistral Medium 3.5DeepSeek-V4-ProvsKimi K3gpt-oss-120bvsKimi K3

Frequently asked questions

DeepSeek-V4-Pro leads gpt-oss-120b on 2 of the 2 benchmarks they both report (BullshitBench v2, GPQA Diamond). gpt-oss-120b shipped 262 days before DeepSeek-V4-Pro, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, DeepSeek-V4-Pro leads at 14% vs gpt-oss-120b at 11%. On GPQA Diamond, DeepSeek-V4-Pro leads at 90.1% vs gpt-oss-120b at 80.1%.