Claude Opus 5vsGrok 4.7

Claude Opus 5
Grok 4.7
Specifications
Context window
1M500k
API pricing
Input price
$5.00$2.00
Output price
$25.00$6.00
Cached input price
$0.50$0.50
Cheapest input
$5.00Amazon Bedrock
Cheapest output
$25.00Amazon Bedrock
Benchmarks
DeepSWE 1.1
68.8%71%
Terminal-Bench 4.0
51.82%38%
HealthBench Professional
59.8%56.7%
Benchmarks
BullshitBench v2
73%
Gray Swan IPI · k = 1
0.2%
Gray Swan IPI · k = 10
1.6%
Gray Swan IPI · k = 15
2%
ProgramBench
4.5%
SWE-Bench Pro
79.2%
SWE-Bench Verified
96%
SWE-Bench Multilingual
89.5%
SWE-Bench Multimodal
59.4%
FrontierCode v1.1 (Main) · main split
53.4%
Next.js Evals
94%
Supabase Evals · with skills
92.8%
Supabase Evals · no skills
89.9%
Frontier-Bench v0.1
43.3%
Terminal-Bench-Science 0.1
29%
BrowseComp
90.8%
Humanity's Last Exam · no tools
56.3%
Humanity's Last Exam · with tools
64.7%
ARC-AGI-3
30.2%
ARC-AGI-2
90.4%
BioMysteryBench · hard
49.4%
BioMysteryBench · human solved
90.1%
EEBench
64%
OSWorld 2.0
70.6%
AutomationBench
26%
Harvey's Legal Agent Benchmark (Held-out)
11.7%
Harvey's Legal Agent Benchmark
19.6%
GDPval-AA v2.1
1695
GDPval-AA v2
1861
AA-Briefcase v1.1
1657
AA-Briefcase
1685
threejseval
1776
Overview
CompanyAnthropicSpaceXAI
Release dateJul 24 2026Sep 21 2026
AccessProprietaryProprietary

Other comparisons

Claude Opus 5vsGPT-6 AstraGrok 4.7vsGPT-6 AstraClaude Opus 5vsGemini 3.8 FlashGrok 4.7vsGemini 3.8 FlashClaude Opus 5vsMuse Spark 1.3Grok 4.7vsMuse Spark 1.3Claude Opus 5vsDeepSeek-V4.1-FlashGrok 4.7vsDeepSeek-V4.1-FlashClaude Opus 5vsMistral Medium 3.5Grok 4.7vsMistral Medium 3.5Claude Opus 5vsKimi K3Grok 4.7vsKimi K3

Frequently asked questions

Claude Opus 5 leads Grok 4.7 on 2 of the 3 benchmarks they both report (DeepSWE 1.1, Terminal-Bench 4.0, HealthBench Professional). Grok 4.7 is cheaper on both input and output: $2.00 vs $5.00 per million input tokens, and $6.00 vs $25.00 per million output tokens. Figures are base-tier rates. Claude Opus 5 shipped 59 days before Grok 4.7, so benchmark comparisons should account for the intervening progress.

Context windows are 1M (Claude Opus 5) vs 500k (Grok 4.7).

On DeepSWE 1.1, Grok 4.7 leads at 71% vs Claude Opus 5 at 68.8%. On Terminal-Bench 4.0, Claude Opus 5 leads at 51.82% vs Grok 4.7 at 38%. On HealthBench Professional, Claude Opus 5 leads at 59.8% vs Grok 4.7 at 56.7%.