# Gemini 3.7 Flash

Gemini 3.7 Flash is an AI model released by Google on Aug 13 2026. It has a 1M context window. Tracked results include 97% on MRCR v2 (8-needle) (128k average), 62.5% on MRCR v2 (8-needle) (1M pointwise) and 35% on BullshitBench v2.

## Facts

| Field | Value |
| --- | --- |
| Model | Gemini 3.7 Flash |
| Developer | Google |
| Release date | Thursday, Aug 13 2026 |
| Licensing | Closed |
| Context window | 1M |

## API pricing

All rates in USD per 1,000,000 tokens, pay-as-you-go.

| Tier | Input | Cached input | Output |
| --- | --- | --- | --- |
| Standard, through 2026-12-31 | $0.75 | $0.075 | $3.75 |

Output prices include reasoning tokens.
Verified August 18, 2026 against the first-party source: https://ai.google.dev/gemini-api/docs/pricing

## Tracked benchmark scores

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| BullshitBench v2 | 35% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better. |
| ProgramBench | 0% | [ProgramBench](https://programbench.com/), retrieved 2026-09-11 | The AI receives a working program and its documentation, then builds a replacement from scratch without the original source code, internet access or decompilation. The score is the percentage of 200 programs that pass every behavioral test. We record each model's best published mini-SWE-agent result, including higher reasoning efforts where available. Partial test-pass rates and almost-solved programs do not count toward this score. Equal scores share a rank here; the official board also uses partial progress to break ties. Higher is better. |
| DeepSWE 1.1 | 65.3% | Lab | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. |
| FrontierCode v1.1 (Main) (main split) | 43.6% | Lab | A set of very hard, frontier-difficulty coding tasks an AI agent has to complete end to end. The score is the share of tasks in the main split it solves. Higher is better. |
| Terminal-Bench 4.0 | 11.21% | [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), retrieved 2026-09-05 | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Version 4.0 recalibrated how much time, CPU and memory each task gets, removed eight tasks and fixed nineteen, so fewer runs fail for reasons that have nothing to do with the model. Scores are not comparable with earlier versions. Higher is better. |
| Terminal-Bench 3.0 | 14.9% | Lab | command-line task completion (v3.0, much harder task set) |
| Terminal-Bench 2.1 | 85.8% | Lab | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better. |
| Humanity's Last Exam (Verified) | 53.6% | Lab | The re-checked edition of Humanity's Last Exam: the same extremely hard expert questions, minus the ones found to be flawed or wrongly answered. Scores on it run lower than on the original exam, so read the two as separate tests rather than a before-and-after. Higher is better. |
| ARC-AGI-2 | 84.6% | [BenchLM](https://benchlm.ai), retrieved 2026-09-28 | Puzzle-style tests of abstract reasoning and pattern-finding — the kind of thing people find easy but AIs often struggle with. Higher is better. |
| BioMysteryBench (hard) | 43.5% | Lab | Real unsolved-style biology puzzles — the AI has to reason its way to an answer the way a research biologist would. The “hard” split contains the toughest cases. Higher is better. |
| BioMysteryBench (human solved) | 87.1% | Lab | Real biology puzzles that human experts have managed to crack — can the AI reach the same answers? Higher is better. |
| LAB-Bench 2 | 82.1% | Lab | Everyday tasks from a working biology lab — reading protocols, interpreting figures and sequence data, and answering the practical questions a researcher hits at the bench. Higher is better. |
| OSWorld 2.0 | 38.1% | Lab | Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Version 2.0 is a harder, refreshed task set. Higher is better. |
| Agent's Last Exam (pass@1) | 26.3% | Lab | A hard set of desktop and operating-system tasks an AI agent has to finish by looking at the screen and working the machine itself. The score is the share it passes outright — partial credit does not count. Higher is better. |
| AutomationBench | 30.4% | Lab | Tests whether the AI can run real multi-step business workflows — the kind of end-to-end office processes companies want to automate — from start to finish. Higher is better. |
| Harvey's Legal Agent Benchmark | 8.8% | Lab | Harvey's test of whether an AI agent can complete real legal work — drafting and reviewing documents, working with spreadsheets and presentations, and navigating files the way a lawyer's assistant would. Higher is better. |
| AA Intelligence Index | 56 | Lab | Artificial Analysis composite intelligence index across evals |
| GDPval-AA v2 | 1525 | Lab | economically valuable knowledge work (v2, re-based Elo) |
| CharXiv Reasoning | 84.5% | Lab | Can the AI read and reason about complex charts and figures, not just text? Higher is better. |
| GDP.PDF | 34% | Lab | Real professional PDFs — filings, reports, technical documents — with questions an expert in that field would ask. Tests whether the AI reads the page as a document, layout and figures included, rather than as loose text. Higher is better. |
| LVBench | 85.4% | Lab | Can the AI follow a very long video — up to an hour — and answer questions that need details from far apart in it? Higher is better. |
| MRCR v2 (8-needle) (128k average) | 97% | Lab | Tests whether the AI can find specific details buried inside a very long document (around 128k tokens — roughly a long book). Higher is better. |
| MRCR v2 (8-needle) (1M pointwise) | 62.5% | Lab | Tests whether the AI can find specific details buried inside an enormous document (around 1 million tokens — many books). Higher is better. |
| threejseval | 1480 | [threejseval](https://threejseval.com/ranking) | Every model gets the same prompt — "the Eiffel Tower", "a glass fishbowl", "a robot arm picking toys into a box" — and builds a 3D scene in Three.js. Real people then see two scenes side by side, names hidden, and vote for the one they prefer. The votes become a chess-style Elo rating on threejseval.com, averaged across all the prompts. It measures whether the scene looks and moves right to a human eye, not whether the code passes a test. Higher is better. |

## About Gemini 3.7 Flash

Gemini 3.7 Flash, released August 13, 2026, arrived three weeks after Gemini 3.6 Flash and moved the biggest numbers of the whole Flash line in coding. At launch it scored 65.3% on DeepSWE v1.1 for long-horizon software engineering, against 49.0% for 3.6 Flash, and 43.6% on the main split of FrontierCode 1.1 — the best figure in Google's launch comparison, ahead of Claude Sonnet 5 and GPT-5.6 Terra. Web development moved with it, to an Arena.ai WebDev Elo of 1588, and the gains extended past code: 30.4% on AutomationBench for enterprise workflows (up from 17.0%), 34.0% on GDP.PDF document comprehension (up from 22.0%), 8.8% all-pass-rate on Harvey's Legal Agent Benchmark for complex legal workflows, and 97.0% on MRCR v2 128k long-context recall. On the Artificial Analysis Intelligence Index it landed at 56, above Sonnet 5 at 55 and a point below GPT-5.6 Terra and Muse Spark 1.2.

Price was the other half of the pitch. Google launched it at an introductory $0.75 per million input tokens and $3.75 per million output through the end of 2026 — half what 3.6 Flash had cost at its own launch — with the rate reverting to $1.50/$7.50 in January 2027. It shipped with a 1M-token context window and multimodal input covering images, video, audio, and PDFs, across the Gemini API, AI Studio, Antigravity, Android Studio, and Gemini Enterprise, and reached consumers through Spark for AI Pro and Ultra subscribers. The agentic computer-use scores were the exception to the sweep: 38.1% on OSWorld 2.0 and 14.9% on Terminal-Bench 3.0 both trailed GPT-5.6 Terra at release. Three Flash-tier upgrades between May and August, against a Pro flagship untouched since February, made clear which tier Google saw as the centre of its lineup. The fourth followed on September 2, 2026, when Gemini 3.8 Flash took over the line after twenty days.

## Questions and answers

### When was Gemini 3.7 Flash released?

Gemini 3.7 Flash was released by Google on Thursday, Aug 13 2026.

### Who made Gemini 3.7 Flash?

Gemini 3.7 Flash was built by Google. Builds the Gemini family of models through Google DeepMind. Integrates AI across Google products.

### How much does Gemini 3.7 Flash cost?

Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens through the Google API. Cached input is $0.075 per million tokens. Output prices include reasoning tokens. Rates are pay-as-you-go API prices verified against Google's published pricing on August 18, 2026.

### What benchmark scores did Gemini 3.7 Flash get?

Gemini 3.7 Flash reports 24 tracked benchmark scores — BullshitBench v2: 35%; ProgramBench: 0%; DeepSWE 1.1: 65.3%; FrontierCode v1.1 (Main) (main split): 43.6%; Terminal-Bench 4.0: 11.21%; Terminal-Bench 3.0: 14.9%; Terminal-Bench 2.1: 85.8%; Humanity's Last Exam (Verified): 53.6%; ARC-AGI-2: 84.6%; BioMysteryBench (hard): 43.5%; BioMysteryBench (human solved): 87.1%; LAB-Bench 2: 82.1%; OSWorld 2.0: 38.1%; Agent's Last Exam (pass@1): 26.3%; AutomationBench: 30.4%; Harvey's Legal Agent Benchmark: 8.8%; AA Intelligence Index: 56; GDPval-AA v2: 1525; CharXiv Reasoning: 84.5%; GDP.PDF: 34%; LVBench: 85.4%; MRCR v2 (8-needle) (128k average): 97%; MRCR v2 (8-needle) (1M pointwise): 62.5%; threejseval: 1480. Tracked scores may come from lab reports or independent benchmarks; source details accompany the benchmark data. It holds the best score among all models tracked here on MRCR v2 (8-needle) (128k average) and MRCR v2 (8-needle) (1M pointwise).

### What is the context window of Gemini 3.7 Flash?

Gemini 3.7 Flash has a context window of 1M. That is the maximum amount of input plus output the model can hold in a single request.

### Is Gemini 3.7 Flash open source?

No. Gemini 3.7 Flash is a closed model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after Gemini 3.7 Flash?

Google's previous tracked release was Gemini 3.5 Flash Cyber on Jul 21 2026, 23 days earlier. It was followed by Gemini 3.8 Flash on Sep 2 2026.


---

Canonical page: https://aireleasetracker.com/model/google/gemini-3.7-flash
Site index: https://aireleasetracker.com/llms.txt
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
