# DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash-0731 is an AI model released by DeepSeek on Jul 31 2026. At release it scored 39% on BullshitBench v2, 79% on SWE-Bench Verified and 34.8% on Humanity's Last Exam (with tools).

## Facts

| Field | Value |
| --- | --- |
| Model | DeepSeek-V4-Flash-0731 |
| Developer | DeepSeek |
| Release date | Friday, Jul 31 2026 |
| Licensing | Proprietary |

## API pricing

All rates in USD per 1,000,000 tokens, pay-as-you-go.

| Tier | Input | Cached input | Output |
| --- | --- | --- | --- |
| Off-peak (all hours outside the published peak windows) | $0.22 | $0.007 | $0.66 |
| Peak (01:00–04:00 and 06:00–10:00 UTC) | $0.44 | $0.014 | $1.32 |

Verified August 18, 2026 against the first-party source: https://api-docs.deepseek.com/quick_start/pricing

## Benchmark scores published at release

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| BullshitBench v2 | 39% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better. |
| SWE-Bench Verified | 79% | [BenchLM](https://benchlm.ai), retrieved 2026-08-31 | Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better. |
| Terminal-Bench 2.1 | 82.7% | Lab | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better. |
| Toolathlon-Verified | 70.3% | Lab | Tests how well the AI uses everyday personal tools and apps to get things done — a human-checked version of Toolathlon. Higher is better. |
| BrowseComp | 73.2% | [BenchLM](https://benchlm.ai), retrieved 2026-08-31 | Can the AI browse the web and track down hard-to-find answers? Higher is better. |
| CyberGym | 76.7% | Lab | Tests the AI on cybersecurity challenges — finding and exploiting software weaknesses inside a safe sandbox. Higher is better. |
| Humanity's Last Exam (with tools) | 34.8% | [BenchLM](https://benchlm.ai), retrieved 2026-08-31 | Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better. |
| AutomationBench | 25.1% | Lab | Tests whether the AI can run real multi-step business workflows — the kind of end-to-end office processes companies want to automate — from start to finish. Higher is better. |

## About DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash-0731, announced July 31, 2026 as the "official" V4-Flash API in public beta, was a post-training refresh of the April V4-Flash build — same architecture, retrained for agentic work. The jump was unusually large for a checkpoint update: on DeepSeek's own evaluation harness it scored 82.7 on Terminal-Bench 2.1 against 61.8 for the April build, and 76.7 on CyberGym, surpassing the larger V4-Pro-Preview on several coding benchmarks despite being the cheaper tier. DeepSeek's harness numbers for competitor models differed from those labs' own published figures, so its launch chart is best read as within-family comparison.

The release doubled as a renaming: DeepSeek retroactively designated the April 24 builds V4-Flash-Preview and V4-Pro-Preview, and served the new checkpoint under the unchanged deepseek-v4-flash API name — existing callers were switched to it without a code change. It added native support for the OpenAI Responses API format and Codex compatibility. Unlike every prior DeepSeek release, it shipped API-only at launch: no 0731 weights were published, and Hugging Face still carried the April build at release. V4-Pro, the app, and the web product stayed on their April checkpoints until the 0813 Pro refresh two weeks later. The Flash tier moved off the V4 architecture entirely with DeepSeek-V4.1-Flash in September 2026.

## Questions and answers

### When was DeepSeek-V4-Flash-0731 released?

DeepSeek-V4-Flash-0731 was released by DeepSeek on Friday, Jul 31 2026.

### Who made DeepSeek-V4-Flash-0731?

DeepSeek-V4-Flash-0731 was built by DeepSeek. Chinese AI lab known for efficient, open-weight models. Gained attention for strong performance at lower cost.

### How much does DeepSeek-V4-Flash-0731 cost?

DeepSeek-V4-Flash-0731 costs $0.22 per million input tokens and $0.66 per million output tokens through the DeepSeek API. Cached input is $0.007 per million tokens. Those are the rates for the “Off-peak (all hours outside the published peak windows)” tier; 1 other pricing tier is published for this model. Rates are pay-as-you-go API prices verified against DeepSeek's published pricing on August 18, 2026.

### What benchmark scores did DeepSeek-V4-Flash-0731 get?

DeepSeek-V4-Flash-0731 reports 8 tracked benchmark scores — BullshitBench v2: 39%; SWE-Bench Verified: 79%; Terminal-Bench 2.1: 82.7%; Toolathlon-Verified: 70.3%; BrowseComp: 73.2%; CyberGym: 76.7%; Humanity's Last Exam (with tools): 34.8%; AutomationBench: 25.1%. Scores are the figures published at release by DeepSeek.

### Is DeepSeek-V4-Flash-0731 open source?

No. DeepSeek-V4-Flash-0731 is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after DeepSeek-V4-Flash-0731?

DeepSeek's previous tracked release was DeepSeek-V4-Flash on Apr 24 2026, 98 days earlier. It was followed by DeepSeek-V4-Pro-0813 on Aug 13 2026.


---

Canonical page: https://aireleasetracker.com/model/deepseek/deepseek-v4-flash-0731
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
