Gemini 3.8 Flash

Released

Gemini 3.8 Flash is an AI model released by Google on Wednesday, Sep 2 2026, 20 days after Gemini 3.7 Flash. It has a 1M token context window. Benchmark results (shown below) cover Terminal-Bench 2.1, Humanity's Last Exam (Verified), BioMysteryBench, LAB-Bench 2, Finance Agent v2, GDP.PDF, and 7 more.

Benchmarks

Coding

DeepSWE 1.1
71%
#2 of 23

Terminal & CLI

Terminal-Bench 2.1
89.4%
#1
Terminal-Bench 4.0
19.1%
#10 of 13

Agentic & tool use

OSWorld 2.0
59%
#2 of 4

Reasoning & science

Humanity's Last Exam (Verified)
54.9%
#1
BioMysteryBench
56.5%
hard
88.8%
human solved
hard
#1
human solved
#2 of 3
LAB-Bench 2
86.2%
#1

Knowledge work

GDPval-AA v2
1545
#10 of 19

Finance

Finance Agent v2
61.4%
#1

Multimodal

GDP.PDF
35%
#1
LVBench
87.1%
87.8%
agentic
#1
agentic
CharXiv Reasoning
86.2%
#4 of 16

About

Gemini 3.8 Flash, released September 2, 2026, was the fourth upgrade to Google's workhorse Flash tier in four months and the one aimed squarely at coding, the area where the line had trailed Anthropic and OpenAI. At launch it scored 89.4% on Terminal-Bench 2.1 for agentic terminal work — the best figure in Google's launch comparison, ahead of Claude Opus 5 at 89.1% and GPT-5.6 Sol at 88.8% — and 71.0% on DeepSWE v1.1 for long-horizon software engineering, up from 65.3% for 3.7 Flash. The harder agentic sets moved with it without closing the gap: 19.1% on Terminal-Bench 4.0 against 51.8% for Opus 5, and 59.0% on OSWorld 2.0 against 75.4%. Google kept the introductory price its predecessor had launched at, $0.75 per million input tokens and $3.75 per million output, against a regular rate of $1.50/$7.50, and the 1M-token context window came across unchanged.

The clearer wins at release were outside code. It led Google's comparison table on financial analyst work (61.4% on Vals Finance Agent v2), on legal workflows (10.0% all-pass-rate on Harvey's Legal Agent Benchmark, against 6.7% for Opus 5), on chart reasoning without tools (86.2% on CharXiv Reasoning), on long video (87.8% on LVBench navigating agentically, 87.1% static), and across the science evaluations — 56.5% on the human-difficult split of BioMysteryBench, seven points clear of the next model, and 86.2% on LAB-Bench 2. Expert reasoning landed at 54.9% on the verified edition of Humanity's Last Exam, within half a point of the frontier models it was priced at a fraction of, and knowledge work at a GDPval-AA v2 Elo of 1545. Read together the release was a Flash model buying most of a frontier model's breadth at Flash prices, while conceding the long-horizon agentic tasks to the tier above it.

Compare Gemini 3.8 Flash with

Gemini 3.8 Flash

Suggested comparisons

Frequently asked questions

Gemini 3.8 Flash was released by Google on Wednesday, Sep 2 2026.

Gemini 3.8 Flash was built by Google. Builds the Gemini family of models through Google DeepMind. Integrates AI across Google products.

Gemini 3.8 Flash reports 15 tracked benchmark scores — DeepSWE 1.1: 71%; Terminal-Bench 4.0: 19.1%; Terminal-Bench 2.1: 89.4%; Humanity's Last Exam (Verified): 54.9%; BioMysteryBench (hard): 56.5%; BioMysteryBench (human solved): 88.8%; LAB-Bench 2: 86.2%; OSWorld 2.0: 59%; Finance Agent v2: 61.4%; Harvey's Legal Agent Benchmark: 10%; GDPval-AA v2: 1545; CharXiv Reasoning: 86.2%; GDP.PDF: 35%; LVBench: 87.1%; LVBench (agentic): 87.8%. Scores are the figures published at release by Google. It holds the best score among all models tracked here on Terminal-Bench 2.1, Humanity's Last Exam (Verified), BioMysteryBench (hard), LAB-Bench 2, Finance Agent v2, GDP.PDF, LVBench and LVBench (agentic).

Gemini 3.8 Flash has a context window of 1M. That is the maximum amount of input plus output the model can hold in a single request.

No. Gemini 3.8 Flash is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

Google's previous tracked release was Gemini 3.7 Flash on Aug 13 2026, 20 days earlier. It was followed by Gemini 3.8 Flash Cyber on Sep 2 2026.

All Google releases

30 tracked