# Muse Spark

Muse Spark is an AI model released by Meta on Apr 8 2026. At release it scored 88.9% on CharXiv Reasoning, 89.5% on GPQA Diamond and 77.4% on SWE-Bench Verified.

## Facts

| Field | Value |
| --- | --- |
| Model | Muse Spark |
| Developer | Meta |
| Release date | Wednesday, Apr 8 2026 |
| Licensing | Proprietary |

## Benchmark scores published at release

| Benchmark | Score | What it measures |
| --- | --- | --- |
| Gray Swan IPI (k = 1) | 2.9% | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. Gray Swan's indirect prompt injection benchmark measures how often such an attack succeeds when the attacker gets a single try. Lower is better. |
| Gray Swan IPI (k = 10) | 14.3% | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. This variant gives the attacker 10 tries and counts an attack as successful if any of them works. Lower is better. |
| Gray Swan IPI (k = 15) | 16.5% | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. This variant gives the attacker 15 tries and counts an attack as successful if any of them works. Lower is better. |
| SWE-Bench Pro | 55% | Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better. |
| SWE-Bench Verified | 77.4% | Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better. |
| DeepSWE 1.1 | 10% | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. |
| Terminal-Bench 2.1 | 67.3% | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better. |
| MCP Atlas | 82.2% | Can the AI chain together many tools and steps to complete one bigger task, rather than doing just a single thing? Higher is better. |
| JobBench | 17% | Tests the AI on professional workplace tasks that require using real work tools — the kind of multi-step jobs an office worker handles. Higher is better. |
| Toolathlon-Verified | 49.4% | Tests how well the AI uses everyday personal tools and apps to get things done — a human-checked version of Toolathlon. Higher is better. |
| Humanity's Last Exam (with tools) | 50.4% | Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better. |
| ARC-AGI-2 | 42.5% | Puzzle-style tests of abstract reasoning and pattern-finding — the kind of thing people find easy but AIs often struggle with. Higher is better. |
| GPQA Diamond | 89.5% | Graduate-level science questions in biology, physics, and chemistry — hard enough that subject-matter PhDs score around 65%. Higher is better. |
| OSWorld-Verified | 53.3% | Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Higher is better. |
| CharXiv Reasoning | 88.9% | Can the AI read and reason about complex charts and figures, not just text? Higher is better. |
| BabyVision | 39.9% | Tests core visual reasoning — seeing and understanding images the way even young children can, which AIs often find surprisingly hard. Higher is better. |
| MMMU | 80.4% | Tests the AI on understanding images and text together across many college subjects. Higher is better. |
| Arena Elo (Text) | 1488 | Real people chat with two anonymous AIs side by side and vote for the answer they prefer. Votes become a chess-style Elo rating on arena.ai — it measures which AI people actually like, not test scores. Higher is better. |

## About Muse Spark

Muse Spark, released April 8, 2026, was Meta's reset: a new model family retiring the Llama name a year after Llama 4, and — a sharp break with Meta's open-weights tradition — released as a proprietary model rather than a downloadable one. It debuted at 89.5% on GPQA Diamond, 80.4% on MMMU, and 77.4% on SWE-Bench Verified.

Its standout results were agentic: 82.2% on MCP Atlas for tool orchestration and 88.9% on CharXiv Reasoning, the best chart-understanding score of any model at release. Muse Spark 1.1 followed on July 9, 2026 with large gains across agentic benchmarks, including 20.0% on Harvey Legal Agent — the top published legal-work score on this tracker, and the coding-focused Muse Spark 1.2 arrived a month later.

## Questions and answers

### When was Muse Spark released?

Muse Spark was released by Meta on Wednesday, Apr 8 2026.

### Who made Muse Spark?

Muse Spark was built by Meta. Develops the open-weight Llama series of models. Committed to open-source AI research.

### What benchmark scores did Muse Spark get?

Muse Spark reports 18 tracked benchmark scores — Gray Swan IPI (k = 1): 2.9%; Gray Swan IPI (k = 10): 14.3%; Gray Swan IPI (k = 15): 16.5%; SWE-Bench Pro: 55%; SWE-Bench Verified: 77.4%; DeepSWE 1.1: 10%; Terminal-Bench 2.1: 67.3%; MCP Atlas: 82.2%; JobBench: 17%; Toolathlon-Verified: 49.4%; Humanity's Last Exam (with tools): 50.4%; ARC-AGI-2: 42.5%; GPQA Diamond: 89.5%; OSWorld-Verified: 53.3%; CharXiv Reasoning: 88.9%; BabyVision: 39.9%; MMMU: 80.4%; Arena Elo (Text): 1488. Scores are the figures published at release by Meta. It holds the best score among all models tracked here on CharXiv Reasoning.

### Is Muse Spark open source?

No. Muse Spark is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after Muse Spark?

Meta's previous tracked release was LLaMA 4 Maverick on Apr 5 2025, 368 days earlier. It was followed by Muse Spark 1.1 on Jul 9 2026.


---

Canonical page: https://aireleasetracker.com/model/meta/muse-spark
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
