# Muse Spark

Muse Spark is an AI model released by Meta on Apr 8 2026. Tracked results include 88.9% on CharXiv Reasoning, 89.5% on GPQA Diamond and 77.4% on SWE-Bench Verified.

## Facts

| Field | Value |
| --- | --- |
| Model | Muse Spark |
| Developer | Meta |
| Release date | Wednesday, Apr 8 2026 |
| Licensing | Closed |

## Tracked benchmark scores

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| Gray Swan IPI (k = 1) | 2.9% | Lab | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. Gray Swan's indirect prompt injection benchmark measures how often such an attack succeeds when the attacker gets a single try. Lower is better. |
| Gray Swan IPI (k = 10) | 14.3% | Lab | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. This variant gives the attacker 10 tries and counts an attack as successful if any of them works. Lower is better. |
| Gray Swan IPI (k = 15) | 16.5% | Lab | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. This variant gives the attacker 15 tries and counts an attack as successful if any of them works. Lower is better. |
| SWE-Bench Pro | 55% | Lab | Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better. |
| SWE-Bench Verified | 77.4% | Lab | Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better. |
| DeepSWE 1.1 | 10% | Lab | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. |
| Terminal-Bench 2.1 | 67.3% | Lab | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better. |
| MCP Atlas | 82.2% | Lab | Can the AI chain together many tools and steps to complete one bigger task, rather than doing just a single thing? Higher is better. |
| JobBench | 17% | Lab | Tests the AI on professional workplace tasks that require using real work tools — the kind of multi-step jobs an office worker handles. Higher is better. |
| Toolathlon-Verified | 49.4% | Lab | Tests how well the AI uses everyday personal tools and apps to get things done — a human-checked version of Toolathlon. Higher is better. |
| Humanity's Last Exam (with tools) | 50.4% | Lab | Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better. |
| ARC-AGI-2 | 42.5% | [BenchLM](https://benchlm.ai), retrieved 2026-07-11 | Puzzle-style tests of abstract reasoning and pattern-finding — the kind of thing people find easy but AIs often struggle with. Higher is better. |
| GPQA Diamond | 89.5% | Lab | Graduate-level science questions in biology, physics, and chemistry — hard enough that subject-matter PhDs score around 65%. Higher is better. |
| OSWorld-Verified | 53.3% | Lab | Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Higher is better. |
| CharXiv Reasoning | 88.9% | Lab | Can the AI read and reason about complex charts and figures, not just text? Higher is better. |
| BabyVision | 39.9% | Lab | Tests core visual reasoning — seeing and understanding images the way even young children can, which AIs often find surprisingly hard. Higher is better. |
| MMMU | 80.4% | Lab | Tests the AI on understanding images and text together across many college subjects. Higher is better. |

## About Muse Spark

Muse Spark, released April 8, 2026, was Meta's reset: a new model family retiring the Llama name a year after Llama 4, and — a sharp break with Meta's open-weights tradition — released as a closed model rather than a downloadable one. It debuted at 89.5% on GPQA Diamond, 80.4% on MMMU, and 77.4% on SWE-Bench Verified.

Its standout results were agentic: 82.2% on MCP Atlas for tool orchestration and 88.9% on CharXiv Reasoning, the best chart-understanding score of any model at release. Muse Spark 1.1 followed on July 9, 2026 with large gains across agentic benchmarks, including 20.0% on Harvey Legal Agent — the top published legal-work score on this tracker, and the coding-focused Muse Spark 1.2 arrived a month later.

## Questions and answers

### When was Muse Spark released?

Muse Spark was released by Meta on Wednesday, Apr 8 2026.

### Who made Muse Spark?

Muse Spark was built by Meta. Develops the open-weight Llama series of models. Committed to open-source AI research.

### What benchmark scores did Muse Spark get?

Muse Spark reports 17 tracked benchmark scores — Gray Swan IPI (k = 1): 2.9%; Gray Swan IPI (k = 10): 14.3%; Gray Swan IPI (k = 15): 16.5%; SWE-Bench Pro: 55%; SWE-Bench Verified: 77.4%; DeepSWE 1.1: 10%; Terminal-Bench 2.1: 67.3%; MCP Atlas: 82.2%; JobBench: 17%; Toolathlon-Verified: 49.4%; Humanity's Last Exam (with tools): 50.4%; ARC-AGI-2: 42.5%; GPQA Diamond: 89.5%; OSWorld-Verified: 53.3%; CharXiv Reasoning: 88.9%; BabyVision: 39.9%; MMMU: 80.4%. Tracked scores may come from lab reports or independent benchmarks; source details accompany the benchmark data. It holds the best score among all models tracked here on CharXiv Reasoning.

### Is Muse Spark open source?

No. Muse Spark is a closed model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after Muse Spark?

Meta's previous tracked release was LLaMA 4 Maverick on Apr 5 2025, 368 days earlier. It was followed by Muse Spark 1.1 on Jul 9 2026.


---

Canonical page: https://aireleasetracker.com/model/meta/muse-spark
Site index: https://aireleasetracker.com/llms.txt
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
