Claude Sonnet 5vsGrok 4.6

Claude Sonnet 5
Grok 4.6
API pricing
Input price
$2.00$2.00
Output price
$10.00$6.00
Cached input price
$0.20$0.50
Cheapest input
$2.00Amazon Bedrock$2.20Amazon Bedrock
Cheapest output
$10.00Amazon Bedrock$6.60Amazon Bedrock
Benchmarks
BullshitBench v2
80%66%
Next.js Evals
81%71%
Terminal-Bench 4.0
12.42%20.3%
threejseval
12731528
Benchmarks
Gray Swan IPI · k = 1
0.6%
Gray Swan IPI · k = 10
4.7%
Gray Swan IPI · k = 15
5.9%
SWE-Bench Pro
63.2%
SWE-Bench Verified
85.2%
DeepSWE 1.1
65.9%
FrontierCode v1.1 (Extended) · extended split
61.3%
APEX-SWE
56.4%
Supabase Evals · with skills
91.3%
Supabase Evals · no skills
79.7%
Frontier-Bench v0.1
14.6%
Terminal-Bench 3.0
26%
Terminal-Bench 2.1
80.4%
APEX-Agents
57.5%
BrowseComp
84.7%
Humanity's Last Exam · no tools
43.2%
Humanity's Last Exam · with tools
57.4%
OSWorld-Verified
81.2%
Harvey's Legal Agent Benchmark
15.8%
AA Intelligence Index
61
GDPval-AA
1618
GDPval-AA v2
1753
AA-Briefcase
1577
Overview
CompanyAnthropicSpaceXAI
Release dateJun 30 2026Aug 12 2026
AccessProprietaryProprietary

Other comparisons

Claude Sonnet 5vsGPT-6 AstraGrok 4.6vsGPT-6 AstraClaude Sonnet 5vsGemini 3.8 FlashGrok 4.6vsGemini 3.8 FlashClaude Sonnet 5vsMuse Spark 1.3Grok 4.6vsMuse Spark 1.3Claude Sonnet 5vsDeepSeek-V4.1-FlashGrok 4.6vsDeepSeek-V4.1-FlashClaude Sonnet 5vsMistral Medium 3.5Grok 4.6vsMistral Medium 3.5Claude Sonnet 5vsKimi K3Grok 4.6vsKimi K3

Frequently asked questions

Claude Sonnet 5 and Grok 4.6 are evenly matched across the 4 benchmarks they both report (BullshitBench v2, Next.js Evals, Terminal-Bench 4.0, threejseval). Both charge $2.00 per million input tokens. Grok 4.6 is cheaper on output: $6.00 vs $10.00 per million tokens. Figures are base-tier rates. Claude Sonnet 5 shipped 43 days before Grok 4.6, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On BullshitBench v2, Claude Sonnet 5 leads at 80% vs Grok 4.6 at 66%. On Next.js Evals, Claude Sonnet 5 leads at 81% vs Grok 4.6 at 71%. On Terminal-Bench 4.0, Grok 4.6 leads at 20.3% vs Claude Sonnet 5 at 12.42%. On threejseval, Grok 4.6 leads at 1528 vs Claude Sonnet 5 at 1273.