Muse Spark

Released

Muse Spark is an AI model released by Meta on Wednesday, Apr 8 2026, 368 days after LLaMA 4 Maverick. Benchmark results (shown below) cover CharXiv Reasoning, Gray Swan IPI, SWE-Bench Pro, SWE-Bench Verified, DeepSWE 1.1, Terminal-Bench 2.1, and 10 more.

Benchmarks

Coding

SWE-Bench Pro
55%
#13 of 20
SWE-Bench Verified
77.4%
#12 of 39
DeepSWE 1.1
10%
#18 of 18

Terminal & CLI

Terminal-Bench 2.1
67.3%
#21 of 25

Agentic & tool use

MCP Atlas
82.2%
#4 of 11
JobBench
17%
#5 of 5
Toolathlon-Verified
49.4%
#5 of 5
OSWorld-Verified
53.3%
#17 of 17

Reasoning & science

Humanity's Last Exam
50.4%
with tools
#17 of 21
ARC-AGI-2
42.5%
#9 of 11
GPQA Diamond
89.5%
#16 of 50

Multimodal

CharXiv Reasoning
88.9%
#1
BabyVision
39.9%
#3 of 3
MMMU
80.4%
#4 of 7

Robustness

Gray Swan IPI
2.9%
k = 1
14.3%
k = 10
16.5%
k = 15
k = 1
#5 of 13
k = 10
#5 of 13
k = 15
#5 of 13

Community preference

Arena Elo (Text)
1488
#9 of 32

About

Muse Spark, released April 8, 2026, was Meta's reset: a new model family retiring the Llama name a year after Llama 4, and — a sharp break with Meta's open-weights tradition — released as a proprietary model rather than a downloadable one. It debuted at 89.5% on GPQA Diamond, 80.4% on MMMU, and 77.4% on SWE-Bench Verified.

Its standout results were agentic: 82.2% on MCP Atlas for tool orchestration and 88.9% on CharXiv Reasoning, the best chart-understanding score of any model at release. Muse Spark 1.1 followed on July 9, 2026 with large gains across agentic benchmarks, including 20.0% on Harvey Legal Agent — the top published legal-work score on this tracker, and the coding-focused Muse Spark 1.2 arrived a month later.

Compare Muse Spark with

Muse Spark

Suggested comparisons

Frequently asked questions

Muse Spark was released by Meta on Wednesday, Apr 8 2026.

Muse Spark was built by Meta. Develops the open-weight Llama series of models. Committed to open-source AI research.

Muse Spark reports 18 tracked benchmark scores — Gray Swan IPI (k = 1): 2.9%; Gray Swan IPI (k = 10): 14.3%; Gray Swan IPI (k = 15): 16.5%; SWE-Bench Pro: 55%; SWE-Bench Verified: 77.4%; DeepSWE 1.1: 10%; Terminal-Bench 2.1: 67.3%; MCP Atlas: 82.2%; JobBench: 17%; Toolathlon-Verified: 49.4%; Humanity's Last Exam (with tools): 50.4%; ARC-AGI-2: 42.5%; GPQA Diamond: 89.5%; OSWorld-Verified: 53.3%; CharXiv Reasoning: 88.9%; BabyVision: 39.9%; MMMU: 80.4%; Arena Elo (Text): 1488. Scores are the figures published at release by Meta. It holds the best score among all models tracked here on CharXiv Reasoning.

No. Muse Spark is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

Meta's previous tracked release was LLaMA 4 Maverick on Apr 5 2025, 368 days earlier. It was followed by Muse Spark 1.1 on Jul 9 2026.

All Meta releases

14 tracked