# Gemini 3.8 Flash

Gemini 3.8 Flash is an AI model released by Google on Sep 2 2026. It has a 1M context window. At release it scored 89.4% on Terminal-Bench 2.1, 54.9% on Humanity's Last Exam (Verified) and 56.5% on BioMysteryBench (hard).

## Facts

| Field | Value |
| --- | --- |
| Model | Gemini 3.8 Flash |
| Developer | Google |
| Release date | Wednesday, Sep 2 2026 |
| Licensing | Proprietary |
| Context window | 1M |

## Benchmark scores published at release

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| DeepSWE 1.1 | 71% | Lab | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. |
| Terminal-Bench 4.0 | 19.1% | Lab | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Version 4.0 recalibrated how much time, CPU and memory each task gets, removed eight tasks and fixed nineteen, so fewer runs fail for reasons that have nothing to do with the model. Scores are not comparable with earlier versions. Higher is better. |
| Terminal-Bench 2.1 | 89.4% | Lab | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better. |
| Humanity's Last Exam (Verified) | 54.9% | Lab | The re-checked edition of Humanity's Last Exam: the same extremely hard expert questions, minus the ones found to be flawed or wrongly answered. Scores on it run lower than on the original exam, so read the two as separate tests rather than a before-and-after. Higher is better. |
| BioMysteryBench (hard) | 56.5% | Lab | Real unsolved-style biology puzzles — the AI has to reason its way to an answer the way a research biologist would. The “hard” split contains the toughest cases. Higher is better. |
| BioMysteryBench (human solved) | 88.8% | Lab | Real biology puzzles that human experts have managed to crack — can the AI reach the same answers? Higher is better. |
| LAB-Bench 2 | 86.2% | Lab | Everyday tasks from a working biology lab — reading protocols, interpreting figures and sequence data, and answering the practical questions a researcher hits at the bench. Higher is better. |
| OSWorld 2.0 | 59% | Lab | Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Version 2.0 is a harder, refreshed task set. Higher is better. |
| Finance Agent v2 | 61.4% | Lab | Tests the AI on real financial-analysis work, like digging through reports and making sound decisions. Higher is better. |
| Harvey's Legal Agent Benchmark | 10% | Lab | Harvey's test of whether an AI agent can complete real legal work — drafting and reviewing documents, working with spreadsheets and presentations, and navigating files the way a lawyer's assistant would. Higher is better. |
| GDPval-AA v2 | 1545 | Lab | economically valuable knowledge work (v2, re-based Elo) |
| CharXiv Reasoning | 86.2% | Lab | Can the AI read and reason about complex charts and figures, not just text? Higher is better. |
| GDP.PDF | 35% | Lab | Real professional PDFs — filings, reports, technical documents — with questions an expert in that field would ask. Tests whether the AI reads the page as a document, layout and figures included, rather than as loose text. Higher is better. |
| LVBench | 87.1% | Lab | Can the AI follow a very long video — up to an hour — and answer questions that need details from far apart in it? Higher is better. |
| LVBench (agentic) | 87.8% | Lab | The same long-video test, run with the AI free to navigate the video itself — skipping around, replaying sections — rather than being shown it once. Read it as a separate test from the unaided run above. Higher is better. |

## About Gemini 3.8 Flash

Gemini 3.8 Flash, released September 2, 2026, was the fourth upgrade to Google's workhorse Flash tier in four months and the one aimed squarely at coding, the area where the line had trailed Anthropic and OpenAI. At launch it scored 89.4% on Terminal-Bench 2.1 for agentic terminal work — the best figure in Google's launch comparison, ahead of Claude Opus 5 at 89.1% and GPT-5.6 Sol at 88.8% — and 71.0% on DeepSWE v1.1 for long-horizon software engineering, up from 65.3% for 3.7 Flash. The harder agentic sets moved with it without closing the gap: 19.1% on Terminal-Bench 4.0 against 51.8% for Opus 5, and 59.0% on OSWorld 2.0 against 75.4%. Google kept the introductory price its predecessor had launched at, $0.75 per million input tokens and $3.75 per million output, against a regular rate of $1.50/$7.50, and the 1M-token context window came across unchanged.

The clearer wins at release were outside code. It led Google's comparison table on financial analyst work (61.4% on Vals Finance Agent v2), on legal workflows (10.0% all-pass-rate on Harvey's Legal Agent Benchmark, against 6.7% for Opus 5), on chart reasoning without tools (86.2% on CharXiv Reasoning), on long video (87.8% on LVBench navigating agentically, 87.1% static), and across the science evaluations — 56.5% on the human-difficult split of BioMysteryBench, seven points clear of the next model, and 86.2% on LAB-Bench 2. Expert reasoning landed at 54.9% on the verified edition of Humanity's Last Exam, within half a point of the frontier models it was priced at a fraction of, and knowledge work at a GDPval-AA v2 Elo of 1545. Read together the release was a Flash model buying most of a frontier model's breadth at Flash prices, while conceding the long-horizon agentic tasks to the tier above it.

## Questions and answers

### When was Gemini 3.8 Flash released?

Gemini 3.8 Flash was released by Google on Wednesday, Sep 2 2026.

### Who made Gemini 3.8 Flash?

Gemini 3.8 Flash was built by Google. Builds the Gemini family of models through Google DeepMind. Integrates AI across Google products.

### What benchmark scores did Gemini 3.8 Flash get?

Gemini 3.8 Flash reports 15 tracked benchmark scores — DeepSWE 1.1: 71%; Terminal-Bench 4.0: 19.1%; Terminal-Bench 2.1: 89.4%; Humanity's Last Exam (Verified): 54.9%; BioMysteryBench (hard): 56.5%; BioMysteryBench (human solved): 88.8%; LAB-Bench 2: 86.2%; OSWorld 2.0: 59%; Finance Agent v2: 61.4%; Harvey's Legal Agent Benchmark: 10%; GDPval-AA v2: 1545; CharXiv Reasoning: 86.2%; GDP.PDF: 35%; LVBench: 87.1%; LVBench (agentic): 87.8%. Scores are the figures published at release by Google. It holds the best score among all models tracked here on Terminal-Bench 2.1, Humanity's Last Exam (Verified), BioMysteryBench (hard), LAB-Bench 2, Finance Agent v2, GDP.PDF, LVBench and LVBench (agentic).

### What is the context window of Gemini 3.8 Flash?

Gemini 3.8 Flash has a context window of 1M. That is the maximum amount of input plus output the model can hold in a single request.

### Is Gemini 3.8 Flash open source?

No. Gemini 3.8 Flash is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after Gemini 3.8 Flash?

Google's previous tracked release was Gemini 3.7 Flash on Aug 13 2026, 20 days earlier. It was followed by Gemini 3.8 Flash Cyber on Sep 2 2026.


---

Canonical page: https://aireleasetracker.com/model/google/gemini-3.8-flash
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
