Grok 4.5vsGLM-5.3

Grok 4.5
GLM-5.3
Specifications
Parameters
743B
Benchmarks
Nonsense detection
BullshitBench v2
54%
Prompt injection robustness
Gray Swan IPI · k = 1
13.4%
Prompt injection robustness
Gray Swan IPI · k = 10
54.2%
Prompt injection robustness
Gray Swan IPI · k = 15
60.8%
Agentic coding
SWE-Bench Pro
64.7%
Multilingual coding
SWE-Bench Multilingual
78%
Agentic coding
CursorBench v3.2
66.7%
Agentic coding
DeepSWE 1.1
54%
66.9%Best
Agentic coding
DeepSWE 1.0
62%
Agentic coding
FrontierCode v1.1 (Extended) · extended split
56.6%
Expert software engineering
APEX-SWE
53.6%
Next.js coding
Next.js Evals
83%
Agentic computer work
Frontier-Bench v0.1
17.8%
Agentic terminal coding
Terminal-Bench 3.0
15.7%
28.3%Best
Agentic terminal coding
Terminal-Bench 2.1
83.3%
Expert agentic work
APEX-Agents
47.1%
Cybersecurity
CyberGym
84.5%
Cybersecurity
ExploitBench
54.4%
Cybersecurity
ExploitGym · 6-hour budget
130
Cybersecurity
ExploitGym · 2-hour budget
105
Multidisciplinary reasoning
Humanity's Last Exam · with tools
62.5%
Agentic computer use
Agent's Last Exam
28.5%
Business workflows
AutomationBench
48.2%
Agentic legal work
Harvey's Legal Agent Benchmark
12.92%
Medical admin work
MedScribe
86.88%
Overall intelligence
AA Intelligence Index
56
Knowledge work
GDPval-AA v2
1526
1769Best
Knowledge work
AA-Briefcase
1313
Community preference
Arena Elo (Text)
1468
Community preference (code)
Arena Elo (Code)
1549
Overview
CompanySpaceXAIZ.ai
Release dateJul 8 2026Aug 14 2026
AccessProprietaryProprietary

Which is better: Grok 4.5 or GLM-5.3?

GLM-5.3 leads Grok 4.5 on 3 of the 3 benchmarks they both report (DeepSWE 1.1, Terminal-Bench 3.0, GDPval-AA v2). Grok 4.5 shipped 37 days before GLM-5.3, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On DeepSWE 1.1, GLM-5.3 leads at 66.9% vs Grok 4.5 at 54%. On Terminal-Bench 3.0, GLM-5.3 leads at 28.3% vs Grok 4.5 at 15.7%. On GDPval-AA v2, GLM-5.3 leads at 1769 vs Grok 4.5 at 1526.

Frequently asked questions

Grok 4.5 was released by SpaceXAI on Jul 8 2026.

GLM-5.3 was released by Z.ai on Aug 14 2026.

GLM-5.3 leads on DeepSWE 1.1 — Grok 4.5 54% vs GLM-5.3 66.9%.

Other comparisons