Muse SparkvsGLM-5

Muse Spark
GLM-5
Specifications
Parameters
744B
Benchmarks
Nonsense detection
BullshitBench v2
28%
Prompt injection robustness
Gray Swan IPI · k = 1
2.9%
Prompt injection robustness
Gray Swan IPI · k = 10
14.3%
Prompt injection robustness
Gray Swan IPI · k = 15
16.5%
Agentic coding
SWE-Bench Pro
55%
Coding
SWE-Bench Verified
77.4%
77.8%Best
Multilingual coding
SWE-Bench Multilingual
73.3%
Agentic coding
DeepSWE 1.1
10%
Agentic terminal coding
Terminal-Bench 2.1
67.3%
Agentic terminal coding
Terminal-Bench 2.0
56.2%
Multi-step tool use
MCP Atlas
82.2%
Professional tool use
JobBench
17%
Personal tool use
Toolathlon-Verified
49.4%
Web browsing
BrowseComp
75.9%
Multidisciplinary reasoning
Humanity's Last Exam · with tools
50.4%
50.4%
Abstract reasoning
ARC-AGI-2
42.5%
Science
GPQA Diamond
89.5%Best
86%
Agentic computer use
OSWorld-Verified
53.3%
Chart reasoning
CharXiv Reasoning
88.9%
Visual reasoning
BabyVision
39.9%
Multimodal
MMMU
80.4%
Community preference
Arena Elo (Text)
1488
Community preference (code)
Arena Elo (Code)
1435
Overview
CompanyMetaZ.ai
Release dateApr 8 2026Feb 12 2026
AccessProprietaryOpen Weight

Which is better: Muse Spark or GLM-5?

Muse Spark and GLM-5 are evenly matched across the 3 benchmarks they both report (SWE-Bench Verified, Humanity's Last Exam, GPQA Diamond). GLM-5 shipped 55 days before Muse Spark, so benchmark comparisons should account for the intervening progress.

Muse Spark is proprietary, while GLM-5 is open weight.

On SWE-Bench Verified, GLM-5 leads at 77.8% vs Muse Spark at 77.4%. On Humanity's Last Exam · with tools, both models score 50.4%. On GPQA Diamond, Muse Spark leads at 89.5% vs GLM-5 at 86%.

Frequently asked questions

Muse Spark was released by Meta on Apr 8 2026.

GLM-5 was released by Z.ai on Feb 12 2026.

GLM-5 leads on SWE-Bench Verified — Muse Spark 77.4% vs GLM-5 77.8%.

Muse Spark is a proprietary model released by Meta. GLM-5 is an open weight model released by Z.ai.

Other comparisons