Muse Spark 1.1vsGrok 4.6

Muse Spark 1.1
Grok 4.6
Benchmarks
Agentic coding
SWE-Bench Pro
61.5%
Agentic coding
CursorBench v3.2
69.9%
Agentic coding
DeepSWE 1.1
53.3%
65.9%Best
Agentic coding
FrontierCode v1.1 (Extended) · extended split
61.3%
Expert software engineering
APEX-SWE
56.4%
Agentic terminal coding
Terminal-Bench 3.0
26%
Agentic terminal coding
Terminal-Bench 2.1
80%
Expert agentic work
APEX-Agents
57.5%
Multi-step tool use
MCP Atlas
88.1%
Professional tool use
JobBench
54.7%
Personal tool use
Toolathlon-Verified
75.6%
Multidisciplinary reasoning
Humanity's Last Exam · with tools
62.1%
Agentic computer use
OSWorld-Verified
80.8%
Agentic financial analysis
Finance Agent v2
57.2%
Agentic legal work
Harvey's Legal Agent Benchmark
20%Best
15.8%
Tax questions
TaxEval v2
79.72%
Medical admin work
MedScribe
88.89%
Overall intelligence
AA Intelligence Index
61
Knowledge work
GDPval-AA v2
1753
Knowledge work
AA-Briefcase
1577
Chart reasoning
CharXiv Reasoning
88.4%
Visual reasoning
BabyVision
76.3%
Community preference
Arena Elo (Text)
1490
Community preference (code)
Arena Elo (Code)
1540
Overview
CompanyMetaSpaceXAI
Release dateJul 9 2026Aug 12 2026
AccessProprietaryProprietary

Which is better: Muse Spark 1.1 or Grok 4.6?

Muse Spark 1.1 and Grok 4.6 are evenly matched across the 2 benchmarks they both report (DeepSWE 1.1, Harvey's Legal Agent Benchmark). Muse Spark 1.1 shipped 34 days before Grok 4.6, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On DeepSWE 1.1, Grok 4.6 leads at 65.9% vs Muse Spark 1.1 at 53.3%. On Harvey's Legal Agent Benchmark, Muse Spark 1.1 leads at 20% vs Grok 4.6 at 15.8%.

Frequently asked questions

Muse Spark 1.1 was released by Meta on Jul 9 2026.

Grok 4.6 was released by SpaceXAI on Aug 12 2026.

Grok 4.6 leads on DeepSWE 1.1 — Muse Spark 1.1 53.3% vs Grok 4.6 65.9%.

Other comparisons