# Muse Spark 1.3

Muse Spark 1.3 is an AI model released by Meta on Sep 2 2026. It has a 1M context window. At release it scored 75.4% on DeepSWE 1.1, 59.4% on SWEAtlas CodeBase QnA and 64.9% on JobBench.

## Facts

| Field | Value |
| --- | --- |
| Model | Muse Spark 1.3 |
| Developer | Meta |
| Release date | Wednesday, Sep 2 2026 |
| Licensing | Proprietary |
| Context window | 1M |

## Benchmark scores published at release

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| DeepSWE 1.1 | 75.4% | Lab | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. |
| SWEAtlas CodeBase QnA | 59.4% | Lab | Questions about how an unfamiliar codebase actually works — where something is handled, what a change would touch — answered by reading the repository rather than editing it. Tests understanding rather than patch-writing. Higher is better. |
| Terminal-Bench 2.1 | 88.8% | Lab | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better. |
| JobBench | 64.9% | Lab | Tests the AI on professional workplace tasks that require using real work tools — the kind of multi-step jobs an office worker handles. Higher is better. |
| DeepSearchQA | 89.4% | Lab | Questions that cannot be answered from one page: the AI has to search the web, follow the trail across several sources, and put the pieces together into an answer. Higher is better. |
| Agentic IF Index (Internal) | 57.8% | Lab | Meta's internal measure of whether a model keeps following the instructions it was given while working as an agent — over a long run of tool calls, not just in a single reply. Higher is better. |
| OSWorld 2.0 | 66.9% | Lab | Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Version 2.0 is a harder, refreshed task set. Higher is better. |
| AutomationBench | 49.4% | Lab | Tests whether the AI can run real multi-step business workflows — the kind of end-to-end office processes companies want to automate — from start to finish. Higher is better. |
| GDPval-AA v2 | 1754 | Lab | economically valuable knowledge work (v2, re-based Elo) |
| MRCR (256k-512k) | 98.5% | Lab | Tests whether the AI can find specific details buried inside a very long document, here across inputs of roughly 256k to 512k tokens — several books' worth of text. Higher is better. |
| MRCR (512k-1M) | 98.1% | Lab | The same buried-detail retrieval test run on even longer inputs, from roughly 512k up to a million tokens. Scores usually slip as the document grows, so read it against the shorter span above. Higher is better. |

## About Muse Spark 1.3

Muse Spark 1.3, released September 2, 2026, arrived four weeks after Muse Spark 1.2 and was announced by Mark Zuckerberg on X rather than through a launch post — a same-day rollout into Muse Code and the Meta Model API, pitched as frontier performance "almost too cheap to meter" and as the biggest jump Meta had made on coding and agentic work. It kept the shape of the 1.2 release, proprietary and with a 1M-token context window, but widened the launch comparison from a handful of coding rows to eleven benchmarks spanning knowledge work, agentic tool use, long context and terminal coding. Meta used the same announcement to trail two unshipped releases — an unnamed model and open weights for the Muse Spark line — giving a date for neither.

The comparison table set Muse Spark 1.3 at its max reasoning setting against Muse Spark 1.2 at xhigh, GPT-5.6 Sol and Claude Opus 5. Its clearest gains at launch were in long context and code comprehension: 98.5% and 98.1% on MRCR at the 256k-512k and 512k-1M spans, against 66.3% and 55.5% for 1.2, with Opus 5 reporting no figure at either span, and 59.4% on SWEAtlas CodeBase QnA, ahead of every rival on the table. It took the coding rows too, with 75.4% on DeepSWE v1.1 and 88.8% on Terminal-Bench 2.1, the latter tied with GPT-5.6 Sol. The agentic and knowledge-work rows went the other way: 1754 on GDPval-AA v2, 64.9% on JobBench, 66.9% on OSWorld 2.0 and 49.4% on AutomationBench each sat just behind Claude Opus 5, while 89.4% on DeepSearchQA and 57.8% on Meta's internal Agentic IF Index trailed both rivals. As at the 1.2 launch, the figures for competing models were Meta's own runs rather than each lab's published results, and several of them differ from the numbers those labs reported themselves.

## Questions and answers

### When was Muse Spark 1.3 released?

Muse Spark 1.3 was released by Meta on Wednesday, Sep 2 2026.

### Who made Muse Spark 1.3?

Muse Spark 1.3 was built by Meta. Develops the open-weight Llama series of models. Committed to open-source AI research.

### What benchmark scores did Muse Spark 1.3 get?

Muse Spark 1.3 reports 11 tracked benchmark scores — DeepSWE 1.1: 75.4%; SWEAtlas CodeBase QnA: 59.4%; Terminal-Bench 2.1: 88.8%; JobBench: 64.9%; DeepSearchQA: 89.4%; Agentic IF Index (Internal): 57.8%; OSWorld 2.0: 66.9%; AutomationBench: 49.4%; GDPval-AA v2: 1754; MRCR (256k-512k): 98.5%; MRCR (512k-1M): 98.1%. Scores are the figures published at release by Meta. It holds the best score among all models tracked here on DeepSWE 1.1, SWEAtlas CodeBase QnA, JobBench, DeepSearchQA, Agentic IF Index (Internal), MRCR (256k-512k) and MRCR (512k-1M).

### What is the context window of Muse Spark 1.3?

Muse Spark 1.3 has a context window of 1M. That is the maximum amount of input plus output the model can hold in a single request.

### Is Muse Spark 1.3 open source?

No. Muse Spark 1.3 is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after Muse Spark 1.3?

Meta's previous tracked release was Muse Glimmer on Aug 10 2026, 23 days earlier. It is the most recent Meta model tracked on AI Release Tracker.


---

Canonical page: https://aireleasetracker.com/model/meta/muse-spark-1.3
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
