# Gemini 3.7 Flash

Gemini 3.7 Flash is an AI model released by Google on Aug 13 2026. It has a 1M context window. At release it scored 53.6% on Humanity's Last Exam (Verified), 82.1% on LAB-Bench 2 and 26.3% on Agent's Last Exam.

## Facts

| Field | Value |
| --- | --- |
| Model | Gemini 3.7 Flash |
| Developer | Google |
| Release date | Thursday, Aug 13 2026 |
| Licensing | Proprietary |
| Context window | 1M |

## Benchmark scores published at release

| Benchmark | Score | What it measures |
| --- | --- | --- |
| DeepSWE 1.1 | 65.3% | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. |
| FrontierCode v1.1 (Main) (main split) | 43.6% | A set of very hard, frontier-difficulty coding tasks an AI agent has to complete end to end. The score is the share of tasks in the main split it solves. Higher is better. |
| Terminal-Bench 3.0 | 14.9% | command-line task completion (v3.0, much harder task set) |
| Terminal-Bench 2.1 | 85.8% | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better. |
| Humanity's Last Exam (Verified) | 53.6% | The re-checked edition of Humanity's Last Exam: the same extremely hard expert questions, minus the ones found to be flawed or wrongly answered. Scores on it run lower than on the original exam, so read the two as separate tests rather than a before-and-after. Higher is better. |
| BioMysteryBench (hard) | 43.5% | Real unsolved-style biology puzzles — the AI has to reason its way to an answer the way a research biologist would. The “hard” split contains the toughest cases. Higher is better. |
| BioMysteryBench (human solved) | 87.1% | Real biology puzzles that human experts have managed to crack — can the AI reach the same answers? Higher is better. |
| LAB-Bench 2 | 82.1% | Everyday tasks from a working biology lab — reading protocols, interpreting figures and sequence data, and answering the practical questions a researcher hits at the bench. Higher is better. |
| OSWorld 2.0 | 38.1% | Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Version 2.0 is a harder, refreshed task set. Higher is better. |
| Agent's Last Exam | 26.3% | A hard set of desktop and operating-system tasks an AI agent has to finish by looking at the screen and working the machine itself. The score is the share it passes outright — partial credit does not count. Higher is better. |
| AutomationBench | 30.4% | Tests whether the AI can run real multi-step business workflows — the kind of end-to-end office processes companies want to automate — from start to finish. Higher is better. |
| Harvey's Legal Agent Benchmark | 90.7% | Harvey's test of whether an AI agent can complete real legal work — drafting and reviewing documents, working with spreadsheets and presentations, and navigating files the way a lawyer's assistant would. Higher is better. |
| AA Intelligence Index | 56 | Artificial Analysis composite intelligence index across evals |
| GDPval-AA v2 | 1525 | economically valuable knowledge work (v2, re-based Elo) |
| CharXiv Reasoning | 84.5% | Can the AI read and reason about complex charts and figures, not just text? Higher is better. |
| GDP.PDF | 34% | Real professional PDFs — filings, reports, technical documents — with questions an expert in that field would ask. Tests whether the AI reads the page as a document, layout and figures included, rather than as loose text. Higher is better. |
| LVBench | 85.4% | Can the AI follow a very long video — up to an hour — and answer questions that need details from far apart in it? Higher is better. |
| MRCR v2 (8-needle) (128k average) | 97% | Tests whether the AI can find specific details buried inside a very long document (around 128k tokens — roughly a long book). Higher is better. |
| MRCR v2 (8-needle) (1M pointwise) | 62.5% | Tests whether the AI can find specific details buried inside an enormous document (around 1 million tokens — many books). Higher is better. |
| Arena Elo (Code) | 1588 | Like the text arena, but people vote on which AI writes better code. The votes become a chess-style Elo rating on arena.ai. Higher is better. |

## About Gemini 3.7 Flash

Gemini 3.7 Flash, released August 13, 2026, arrived three weeks after Gemini 3.6 Flash and moved the biggest numbers of the whole Flash line in coding. At launch it scored 65.3% on DeepSWE v1.1 for long-horizon software engineering, against 49.0% for 3.6 Flash, and 43.6% on the main split of FrontierCode 1.1 — the best figure in Google's launch comparison, ahead of Claude Sonnet 5 and GPT-5.6 Terra. Web development moved with it, to an Arena.ai WebDev Elo of 1588, and the gains extended past code: 30.4% on AutomationBench for enterprise workflows (up from 17.0%), 34.0% on GDP.PDF document comprehension (up from 22.0%), 90.7% on Harvey's LAB-AA legal benchmark, and 97.0% on MRCR v2 128k long-context recall. On the Artificial Analysis Intelligence Index it landed at 56, above Sonnet 5 at 55 and a point below GPT-5.6 Terra and Muse Spark 1.2.

Price was the other half of the pitch. Google launched it at an introductory $0.75 per million input tokens and $3.75 per million output through the end of 2026 — half what 3.6 Flash had cost at its own launch — with the rate reverting to $1.50/$7.50 in January 2027. It shipped with a 1M-token context window and multimodal input covering images, video, audio, and PDFs, across the Gemini API, AI Studio, Antigravity, Android Studio, and Gemini Enterprise, and reached consumers through Spark for AI Pro and Ultra subscribers. The agentic computer-use scores were the exception to the sweep: 38.1% on OSWorld 2.0 and 14.9% on Terminal-Bench 3.0 both trailed GPT-5.6 Terra at release. Three Flash-tier upgrades between May and August, against a Pro flagship untouched since February, made clear which tier Google saw as the centre of its lineup.

## Questions and answers

### When was Gemini 3.7 Flash released?

Gemini 3.7 Flash was released by Google on Thursday, Aug 13 2026.

### Who made Gemini 3.7 Flash?

Gemini 3.7 Flash was built by Google. Builds the Gemini family of models through Google DeepMind. Integrates AI across Google products.

### What benchmark scores did Gemini 3.7 Flash get?

Gemini 3.7 Flash reports 20 tracked benchmark scores — DeepSWE 1.1: 65.3%; FrontierCode v1.1 (Main) (main split): 43.6%; Terminal-Bench 3.0: 14.9%; Terminal-Bench 2.1: 85.8%; Humanity's Last Exam (Verified): 53.6%; BioMysteryBench (hard): 43.5%; BioMysteryBench (human solved): 87.1%; LAB-Bench 2: 82.1%; OSWorld 2.0: 38.1%; Agent's Last Exam: 26.3%; AutomationBench: 30.4%; Harvey's Legal Agent Benchmark: 90.7%; AA Intelligence Index: 56; GDPval-AA v2: 1525; CharXiv Reasoning: 84.5%; GDP.PDF: 34%; LVBench: 85.4%; MRCR v2 (8-needle) (128k average): 97%; MRCR v2 (8-needle) (1M pointwise): 62.5%; Arena Elo (Code): 1588. Scores are the figures published at release by Google. It holds the best score among all models tracked here on Humanity's Last Exam (Verified), LAB-Bench 2, Agent's Last Exam, Harvey's Legal Agent Benchmark, GDP.PDF, LVBench, MRCR v2 (8-needle) (128k average) and MRCR v2 (8-needle) (1M pointwise).

### What is the context window of Gemini 3.7 Flash?

Gemini 3.7 Flash has a context window of 1M. That is the maximum amount of input plus output the model can hold in a single request.

### Is Gemini 3.7 Flash open source?

No. Gemini 3.7 Flash is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after Gemini 3.7 Flash?

Google's previous tracked release was Gemini 3.5 Flash Cyber on Jul 21 2026, 23 days earlier. It is the most recent Google model tracked on AI Release Tracker.


---

Canonical page: https://aireleasetracker.com/model/google/gemini-3.7-flash
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
