Claude 3 SonnetvsDeepSeek-V4.1-Flash

Claude 3 Sonnet
DeepSeek-V4.1-Flash
Benchmarks
GPQA Diamond
40.4%90.9%
Benchmarks
DeepSWE 1.1
74.2%
NL2Repo-Bench
65.4%
Terminal-Bench 4.0
31.2%
Terminal-Bench 3.0
30%
Terminal-Bench 2.1
90.6%
CyberGym
88.1%
Humanity's Last Exam · no tools
36.8%
Humanity's Last Exam · with tools
63.9%
Agent's Last Exam · pass@1
31.8%
AutomationBench
54.8%
Chartography · with tools
78.9%
Overview
CompanyAnthropicDeepSeek
Release dateMar 4 2024Sep 10 2026
AccessProprietaryProprietary

Other comparisons

Claude 3 SonnetvsGPT-6 AstraDeepSeek-V4.1-FlashvsGPT-6 AstraClaude 3 SonnetvsGemini 3.8 FlashDeepSeek-V4.1-FlashvsGemini 3.8 FlashClaude 3 SonnetvsMuse Spark 1.3DeepSeek-V4.1-FlashvsMuse Spark 1.3Claude 3 SonnetvsGrok 4.6DeepSeek-V4.1-FlashvsGrok 4.6Claude 3 SonnetvsMistral Medium 3.5DeepSeek-V4.1-FlashvsMistral Medium 3.5Claude 3 SonnetvsKimi K3DeepSeek-V4.1-FlashvsKimi K3

Frequently asked questions

DeepSeek-V4.1-Flash leads Claude 3 Sonnet on 1 of the 1 benchmark they both report (GPQA Diamond). Claude 3 Sonnet shipped 920 days before DeepSeek-V4.1-Flash, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On GPQA Diamond, DeepSeek-V4.1-Flash leads at 90.9% vs Claude 3 Sonnet at 40.4%.