Claude 3.5 SonnetvsDeepSeek-V4.1-Flash

Claude 3.5 Sonnet
DeepSeek-V4.1-Flash
Benchmarks
GPQA Diamond
59.4%90.9%
Benchmarks
BullshitBench v2
45%
SWE-Bench Verified
33.4%
DeepSWE 1.1
74.2%
NL2Repo-Bench
65.4%
Terminal-Bench 4.0
31.2%
Terminal-Bench 3.0
30%
Terminal-Bench 2.1
90.6%
CyberGym
88.1%
Humanity's Last Exam · no tools
36.8%
Humanity's Last Exam · with tools
63.9%
Agent's Last Exam · pass@1
31.8%
AutomationBench
54.8%
Chartography · with tools
78.9%
Overview
CompanyAnthropicDeepSeek
Release dateJun 20 2024Sep 10 2026
AccessProprietaryProprietary

Other comparisons

Claude 3.5 SonnetvsGPT-6 AstraDeepSeek-V4.1-FlashvsGPT-6 AstraClaude 3.5 SonnetvsGemini 3.8 FlashDeepSeek-V4.1-FlashvsGemini 3.8 FlashClaude 3.5 SonnetvsMuse Spark 1.3DeepSeek-V4.1-FlashvsMuse Spark 1.3Claude 3.5 SonnetvsGrok 4.6DeepSeek-V4.1-FlashvsGrok 4.6Claude 3.5 SonnetvsMistral Medium 3.5DeepSeek-V4.1-FlashvsMistral Medium 3.5Claude 3.5 SonnetvsKimi K3DeepSeek-V4.1-FlashvsKimi K3

Frequently asked questions

DeepSeek-V4.1-Flash leads Claude 3.5 Sonnet on 1 of the 1 benchmark they both report (GPQA Diamond). Claude 3.5 Sonnet shipped 812 days before DeepSeek-V4.1-Flash, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On GPQA Diamond, DeepSeek-V4.1-Flash leads at 90.9% vs Claude 3.5 Sonnet at 59.4%.