Claude Sonnet 5vsDeepSeek-V4.1-Flash

Claude Sonnet 5
DeepSeek-V4.1-Flash
API pricing
Input price
$2.00
Output price
$10.00
Cached input price
$0.20
Cheapest input
$2.00Amazon Bedrock
Cheapest output
$10.00Amazon Bedrock
Benchmarks
Terminal-Bench 4.0
12.42%31.2%
Terminal-Bench 2.1
80.4%90.6%
Humanity's Last Exam · no tools
43.2%36.8%
Humanity's Last Exam · with tools
57.4%63.9%
Benchmarks
BullshitBench v2
80%
Gray Swan IPI · k = 1
0.6%
Gray Swan IPI · k = 10
4.7%
Gray Swan IPI · k = 15
5.9%
SWE-Bench Pro
63.2%
SWE-Bench Verified
85.2%
DeepSWE 1.1
74.2%
Next.js Evals
81%
Supabase Evals · with skills
88.4%
Supabase Evals · no skills
81.2%
NL2Repo-Bench
65.4%
Frontier-Bench v0.1
14.6%
Terminal-Bench 3.0
30%
BrowseComp
84.7%
CyberGym
88.1%
GPQA Diamond
90.9%
OSWorld-Verified
81.2%
Agent's Last Exam · pass@1
31.8%
AutomationBench
54.8%
GDPval-AA
1618
Chartography · with tools
78.9%
threejseval
1252
Overview
CompanyAnthropicDeepSeek
Release dateJun 30 2026Sep 10 2026
AccessProprietaryProprietary

Other comparisons

Claude Sonnet 5vsGPT-6 AstraDeepSeek-V4.1-FlashvsGPT-6 AstraClaude Sonnet 5vsGemini 3.8 FlashDeepSeek-V4.1-FlashvsGemini 3.8 FlashClaude Sonnet 5vsMuse Spark 1.3DeepSeek-V4.1-FlashvsMuse Spark 1.3Claude Sonnet 5vsGrok 4.6DeepSeek-V4.1-FlashvsGrok 4.6Claude Sonnet 5vsMistral Medium 3.5DeepSeek-V4.1-FlashvsMistral Medium 3.5Claude Sonnet 5vsKimi K3DeepSeek-V4.1-FlashvsKimi K3

Frequently asked questions

DeepSeek-V4.1-Flash leads Claude Sonnet 5 on 3 of the 4 benchmarks they both report (Terminal-Bench 4.0, Terminal-Bench 2.1, Humanity's Last Exam). Only Claude Sonnet 5 has a verified first-party API price: $2.00 per million input tokens and $10.00 per million output tokens. No pay-as-you-go API rate is tracked for DeepSeek-V4.1-Flash. Claude Sonnet 5 shipped 72 days before DeepSeek-V4.1-Flash, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On Terminal-Bench 4.0, DeepSeek-V4.1-Flash leads at 31.2% vs Claude Sonnet 5 at 12.42%. On Terminal-Bench 2.1, DeepSeek-V4.1-Flash leads at 90.6% vs Claude Sonnet 5 at 80.4%. On Humanity's Last Exam · no tools, Claude Sonnet 5 leads at 43.2% vs DeepSeek-V4.1-Flash at 36.8%. On Humanity's Last Exam · with tools, DeepSeek-V4.1-Flash leads at 63.9% vs Claude Sonnet 5 at 57.4%.