Grok 4.6vsGLM-5.3

Grok 4.6
GLM-5.3
Specifications
Parameters
743B
Benchmarks
Agentic coding
CursorBench v3.2
69.9%
Agentic coding
DeepSWE 1.1
65.9%
66.9%Best
Agentic coding
FrontierCode v1.1 (Extended) · extended split
61.3%
Expert software engineering
APEX-SWE
56.4%
Agentic terminal coding
Terminal-Bench 3.0
26%
28.3%Best
Expert agentic work
APEX-Agents
57.5%
Cybersecurity
CyberGym
84.5%
Cybersecurity
ExploitBench
54.4%
Cybersecurity
ExploitGym · 6-hour budget
130
Cybersecurity
ExploitGym · 2-hour budget
105
Multidisciplinary reasoning
Humanity's Last Exam · with tools
62.5%
Agentic computer use
Agent's Last Exam
28.5%
Business workflows
AutomationBench
48.2%
Agentic legal work
Harvey's Legal Agent Benchmark
15.8%
Overall intelligence
AA Intelligence Index
61
Knowledge work
GDPval-AA v2
1753
1769Best
Knowledge work
AA-Briefcase
1577
Overview
CompanySpaceXAIZ.ai
Release dateAug 12 2026Aug 14 2026
AccessProprietaryProprietary

Which is better: Grok 4.6 or GLM-5.3?

GLM-5.3 leads Grok 4.6 on 3 of the 3 benchmarks they both report (DeepSWE 1.1, Terminal-Bench 3.0, GDPval-AA v2). Grok 4.6 shipped 2 days before GLM-5.3, so benchmark comparisons should account for the intervening progress.

Published specifications for these two models are limited — see each model page for the latest details.

On DeepSWE 1.1, GLM-5.3 leads at 66.9% vs Grok 4.6 at 65.9%. On Terminal-Bench 3.0, GLM-5.3 leads at 28.3% vs Grok 4.6 at 26%. On GDPval-AA v2, GLM-5.3 leads at 1769 vs Grok 4.6 at 1753.

Frequently asked questions

Grok 4.6 was released by SpaceXAI on Aug 12 2026.

GLM-5.3 was released by Z.ai on Aug 14 2026.

GLM-5.3 leads on DeepSWE 1.1 — Grok 4.6 65.9% vs GLM-5.3 66.9%.

Other comparisons