DeepSeek-V3.2vsGrok 4.1
DeepSeek-V3.2 | Grok 4.1 | |
|---|---|---|
| Benchmarks | ||
Arena Elo (Code)Community preference (code) — Like the text arena, but people vote on which AI writes better code. The votes become a chess-style Elo rating on arena.ai. Higher is better. | 1361 | 1209 |
| BenchmarksPublished by one model only | ||
BullshitBench v2Nonsense detection — Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better. | 13% | — |
Arena Elo (Text)Community preference — Real people chat with two anonymous AIs side by side and vote for the answer they prefer. Votes become a chess-style Elo rating on arena.ai — it measures which AI people actually like, not test scores. Higher is better. | — | 1466 |
| Overview | ||
| Company | DeepSeek | SpaceXAI |
| Release date | Dec 1 2025 | Nov 17 2025 |
| Access | Open Weight | Proprietary |
Frequently asked questions
DeepSeek-V3.2 leads Grok 4.1 on 1 of the 1 benchmark they both report (Arena Elo (Code)). Grok 4.1 shipped 14 days before DeepSeek-V3.2, so benchmark comparisons should account for the intervening progress.
DeepSeek-V3.2 is open weight, while Grok 4.1 is proprietary.
On Arena Elo (Code), DeepSeek-V3.2 leads at 1361 vs Grok 4.1 at 1209.
DeepSeek-V3.2 was released by DeepSeek on Dec 1 2025.
Grok 4.1 was released by SpaceXAI on Nov 17 2025.
DeepSeek-V3.2 is an open weight model released by DeepSeek. Grok 4.1 is a proprietary model released by SpaceXAI.
Other comparisons
DeepSeek-V3.2vsClaude Opus 5Grok 4.1vsClaude Opus 5DeepSeek-V3.2vsGPT-5.6-CyberGrok 4.1vsGPT-5.6-CyberDeepSeek-V3.2vsGemini 3.7 FlashGrok 4.1vsGemini 3.7 FlashDeepSeek-V3.2vsMuse GlimmerGrok 4.1vsMuse GlimmerDeepSeek-V3.2vsMistral Medium 3.5Grok 4.1vsMistral Medium 3.5DeepSeek-V3.2vsKimi K3Grok 4.1vsKimi K3