Gemini 4 Argon

Released

Gemini 4 Argon is an AI model released by Google on Wednesday, Sep 30 2026, 28 days after Gemini 3.8 Flash Cyber. Benchmark results (shown below) cover Gray Swan IPI, DeepSWE 1.1, FrontierSWE V2, LAB-Bench 2, Finance Agent v2, LVBench, and 6 more.

API pricing

Input
$2.00
Cached input
$0.10
Output
$10.00
  • Reasoning or thinking is supported.
  • Introductory pricing. Once it expires, $4 per million input tokens and $20 per million output tokens apply.
  • Rolling out first to trusted cyber defenders through the Fairwind Program, then to paid API customers and Google AI Ultra subscribers.
USD per 1M tokens

Benchmarks

Coding

DeepSWE 1.1
77.9%
#1
FrontierSWE V2
55%
#1

Terminal & CLI

Terminal-Bench 4.0
57.4%
#6 of 20
Terminal-Bench-Science 0.1
57.6%
#3 of 7

Agentic & tool use

OSWorld 2.0
69.2%
#5 of 10
Agent's Last Exam
39.5%pass@1
#3 of 9
AutomationBench
51.3%
#2 of 15

Reasoning & science

LAB-Bench 2
88.8%
#1

Finance

Finance Agent v2
65.4%
#1

Multimodal

LVBench
91.7%
#1

Robustness

Gray Swan IPI
0.7%k = 15
#1

Compare Gemini 4 Argon with

Gemini 4 Argon

About

Gemini 4 Argon, announced September 30, 2026, was Google's first new frontier model since Gemini 3.1 Pro in February, after seven months in which every Gemini upgrade had landed on the Flash tier. Google did not open it to everyone at once: Argon went first to trusted cyber defenders through its Fairwind Program, unrestricted by cyber guardrails, with paid API customers and Google AI Ultra subscribers promised next, once early-tester feedback had shaped the safeguards. Google also put it through the US government's voluntary pre-release access process. The announced introductory price was $2 per million input tokens and $10 per million output, with cached input 95% cheaper, rising to $4 and $20 once the introductory period ended, and the output limit rose to 1M tokens from the previous 64K, so that a single trajectory could think and write for hundreds of thousands of tokens.

Google's launch table set it against GPT-6 Astra and led with long-horizon coding and knowledge work. It scored 77.9% on DeepSWE v1.1, which Google called a new state of the art, 51.3% on AutomationBench, 65.4% on Vals Finance Agent v2, 19.6% on Harvey's Legal Agent Benchmark and 68.9% on the Vals Index, all ahead of Astra in Google's comparison, along with 91.7% on LVBench for long video. Google also reported a 0.7% attack success rate on Gray Swan's indirect prompt injection benchmark at fifteen attempts, the lowest in its comparison, against 1.0% for Claude Opus 5.5 and Claude Fable 5.1. Astra came out ahead on FrontierSWE v2 (65.5% against Argon's 55.0%), Terminal-Bench Science 0.1 and OSWorld 2.0, and the two were close on Terminal-Bench 4.0, where Argon scored 57.4%. Security was the other half of the pitch: Argon tied for first on CWE-bench v1 at 68.0%, and in an early run through Wiz's Scan for Good programme it found a critical flaw exposing personal data in healthcare software used by hospitals worldwide. Internally, Google said, Argon agents had freed more than 300 TiB of memory across its data centres and were porting C and C++ code to Rust, from libraries like re2 up to the 800,000-line Fuchsia Zircon kernel.

Frequently asked questions

Gemini 4 Argon was released by Google on Wednesday, Sep 30 2026.

All Google releases

31 tracked