# DeepSeek-V4.1-Flash

DeepSeek-V4.1-Flash is an AI model released by DeepSeek on Sep 10 2026. At release it scored 90.6% on Terminal-Bench 2.1, 65.4% on NL2Repo-Bench and 30% on Terminal-Bench 3.0.

## Facts

| Field | Value |
| --- | --- |
| Model | DeepSeek-V4.1-Flash |
| Developer | DeepSeek |
| Release date | Thursday, Sep 10 2026 |
| Licensing | Proprietary |

## Benchmark scores published at release

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| DeepSWE 1.1 | 74.2% | Lab | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. |
| NL2Repo-Bench | 65.4% | Lab | Tests whether the AI can turn a natural-language requirement into working code across an entire repository, not just produce a single function or patch. Higher is better. |
| Terminal-Bench 4.0 | 31.2% | Lab | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Version 4.0 recalibrated how much time, CPU and memory each task gets, removed eight tasks and fixed nineteen, so fewer runs fail for reasons that have nothing to do with the model. Scores are not comparable with earlier versions. Higher is better. |
| Terminal-Bench 3.0 | 30% | Lab | command-line task completion (v3.0, much harder task set) |
| Terminal-Bench 2.1 | 90.6% | Lab | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better. |
| CyberGym | 88.1% | Lab | Tests the AI on cybersecurity challenges — finding and exploiting software weaknesses inside a safe sandbox. Higher is better. |
| Humanity's Last Exam (no tools) | 36.8% | Lab | Humanity's Last Exam — extremely hard expert questions across many subjects, written so you can't just look up the answer. “No tools” means the AI answers on its own. Higher is better. |
| Humanity's Last Exam (with tools) | 63.9% | Lab | Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better. |
| GPQA Diamond | 90.9% | Lab | Graduate-level science questions in biology, physics, and chemistry — hard enough that subject-matter PhDs score around 65%. Higher is better. |
| Agent's Last Exam (pass@1) | 31.8% | Lab | A hard set of desktop and operating-system tasks an AI agent has to finish by looking at the screen and working the machine itself. The score is the share it passes outright — partial credit does not count. Higher is better. |
| AutomationBench | 54.8% | Lab | Tests whether the AI can run real multi-step business workflows — the kind of end-to-end office processes companies want to automate — from start to finish. Higher is better. |
| Chartography (with tools) | 78.9% | Lab | A chart-centred test run with tools available to the AI, reported separately from the chart-reading benchmarks above it. Higher is better. |

## About DeepSeek-V4.1-Flash

DeepSeek-V4.1-Flash, released September 10, 2026, opened a new architecture family for the lab — and did it from the bottom, as the smallest model in the line rather than a flagship. DeepSeek described it as built for greater capability, faster inference and higher throughput, with the architecture meant to scale up to larger models still to come, and with native visual understanding rather than a vision encoder attached after the fact. On the lab's own harness it scored 90.6 on Terminal-Bench 2.1, 74.2 on DeepSWE v1.1 and 88.1 on CyberGym, all ahead of the larger V4-Pro-0813 from four weeks earlier, and 54.8 on AutomationBench. Terminal-Bench 3.0 at 30.0 and 4.0 at 31.2 were around two and a half times the V4-Pro-0813 figures on the same boards.

The visual side of the launch was the genuinely new part for DeepSeek, whose V3 and V4 lines had been text-only: the launch table reported 78.9 on Chartography and 89.6 on BabyVision, both run with tools, putting a DeepSeek model on multimodal boards for the first time. Reasoning was the softer spot — 36.8% on Humanity's Last Exam without tools, the one unstarred figure on a row where the other DeepSeek columns reported the text-only subset, while 63.9% with tools was the highest in the comparison. As with the July and August snapshots, the code-agent numbers came from DeepSeek's own unreleased harness and its figures for rival models did not always match those labs' published ones, so the chart reads best as a within-family comparison. The launch thread announced no weight release.

## Questions and answers

### When was DeepSeek-V4.1-Flash released?

DeepSeek-V4.1-Flash was released by DeepSeek on Thursday, Sep 10 2026.

### Who made DeepSeek-V4.1-Flash?

DeepSeek-V4.1-Flash was built by DeepSeek. Chinese AI lab known for efficient, open-weight models. Gained attention for strong performance at lower cost.

### What benchmark scores did DeepSeek-V4.1-Flash get?

DeepSeek-V4.1-Flash reports 12 tracked benchmark scores — DeepSWE 1.1: 74.2%; NL2Repo-Bench: 65.4%; Terminal-Bench 4.0: 31.2%; Terminal-Bench 3.0: 30%; Terminal-Bench 2.1: 90.6%; CyberGym: 88.1%; Humanity's Last Exam (no tools): 36.8%; Humanity's Last Exam (with tools): 63.9%; GPQA Diamond: 90.9%; Agent's Last Exam (pass@1): 31.8%; AutomationBench: 54.8%; Chartography (with tools): 78.9%. Scores are the figures published at release by DeepSeek. It holds the best score among all models tracked here on NL2Repo-Bench, Terminal-Bench 3.0, Terminal-Bench 2.1, CyberGym, AutomationBench and Chartography (with tools).

### Is DeepSeek-V4.1-Flash open source?

No. DeepSeek-V4.1-Flash is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after DeepSeek-V4.1-Flash?

DeepSeek's previous tracked release was DeepSeek-V4-Pro-0813 on Aug 13 2026, 28 days earlier. It is the most recent DeepSeek model tracked on AI Release Tracker.


---

Canonical page: https://aireleasetracker.com/model/deepseek/deepseek-v4.1-flash
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
