Compare AI models

Specifications
Parameters
550B117B
Context window
—128k
API pricing
Cheapest input
$0.50DeepInfra$0.03AkashML
Cheapest output
$2.20DeepInfra$0.15Venice
Benchmarks
BullshitBench v2
49%12%
SWE-Bench Verified
71.9%62.4%
Humanity's Last Exam · with tools
26.7%19%
GPQA Diamond
86.7%80.1%
Overview
CompanyNVIDIAOpenAI
Release dateJun 4 2026Aug 5 2025
AccessOpen SourceOpen Weight
Model detailsView modelView model

Frequently asked questions

Nemotron 3 Ultra leads gpt-oss-120b on 4 of the 4 benchmarks they both report (BullshitBench v2, SWE-Bench Verified, Humanity's Last Exam, GPQA Diamond). gpt-oss-120b shipped 303 days before Nemotron 3 Ultra, so benchmark comparisons should account for the intervening progress.

Nemotron 3 Ultra has 550B parameters, while gpt-oss-120b has 117B. Nemotron 3 Ultra is open source, while gpt-oss-120b is open weight.

On BullshitBench v2, Nemotron 3 Ultra leads at 49% vs gpt-oss-120b at 12%. On SWE-Bench Verified, Nemotron 3 Ultra leads at 71.9% vs gpt-oss-120b at 62.4%. On Humanity's Last Exam · with tools, Nemotron 3 Ultra leads at 26.7% vs gpt-oss-120b at 19%. On GPQA Diamond, Nemotron 3 Ultra leads at 86.7% vs gpt-oss-120b at 80.1%.