Grok 4.5vsGLM-5.3-Flash

Grok 4.5
GLM-5.3-Flash
Specifications
Parameters
320B
Context window
1M
API pricing
Input price
$2.00
Output price
$6.00
Cached input price
$0.30
Cheapest input
$0.075DeepInfra
Cheapest output
$0.25DeepInfra
Benchmarks
DeepSWE 1.1
54%63.4%
Terminal-Bench 2.1
83.3%84.3%
GDPval-AA v2
15261773
Benchmarks
BullshitBench v2
54%
Gray Swan IPI · k = 1
13.4%
Gray Swan IPI · k = 10
54.2%
Gray Swan IPI · k = 15
60.8%
SWE-Bench Pro
64.7%
SWE-Bench Multilingual
78%
DeepSWE 1.0
62%
FrontierCode v1.1 (Extended) · extended split
56.6%
APEX-SWE
53.6%
Next.js Evals
77%
NL2Repo-Bench
56.3%
Frontier-Bench v0.1
17.8%
Terminal-Bench 4.0
12.42%
Terminal-Bench 3.0
15.7%
APEX-Agents
47.1%
Toolathlon-Verified
78.4%
Humanity's Last Exam · with tools
55.3%
ARC-AGI-2
52.64%
Agent's Last Exam · pass@1
26.3%
AutomationBench
48.8%
Harvey's Legal Agent Benchmark
12.92%
MedScribe
86.88%
AA Intelligence Index
56
AA-Briefcase
1313
CharXiv Reasoning · with tools
89.4%
Chartography · with tools
78%
OfficeQA Pro
62.4%
MVBench
77.8%
MMVU
80.5%
BabyVision
53.4%
Overview
CompanySpaceXAIZ.ai
Release dateJul 8 2026Aug 26 2026
AccessProprietaryOpen Weight

Other comparisons

Grok 4.5vsClaude Fable 5.1GLM-5.3-FlashvsClaude Fable 5.1Grok 4.5vsGPT-6 AstraGLM-5.3-FlashvsGPT-6 AstraGrok 4.5vsGemini 3.8 FlashGLM-5.3-FlashvsGemini 3.8 FlashGrok 4.5vsMuse Spark 1.3GLM-5.3-FlashvsMuse Spark 1.3Grok 4.5vsDeepSeek-V4-Pro-0813GLM-5.3-FlashvsDeepSeek-V4-Pro-0813Grok 4.5vsMistral Medium 3.5GLM-5.3-FlashvsMistral Medium 3.5

Frequently asked questions

GLM-5.3-Flash leads Grok 4.5 on 3 of the 3 benchmarks they both report (DeepSWE 1.1, Terminal-Bench 2.1, GDPval-AA v2). Only Grok 4.5 has a verified first-party API price: $2.00 per million input tokens and $6.00 per million output tokens. No pay-as-you-go API rate is tracked for GLM-5.3-Flash. Grok 4.5 shipped 49 days before GLM-5.3-Flash, so benchmark comparisons should account for the intervening progress.

Grok 4.5 is proprietary, while GLM-5.3-Flash is open weight.

On DeepSWE 1.1, GLM-5.3-Flash leads at 63.4% vs Grok 4.5 at 54%. On Terminal-Bench 2.1, GLM-5.3-Flash leads at 84.3% vs Grok 4.5 at 83.3%. On GDPval-AA v2, GLM-5.3-Flash leads at 1773 vs Grok 4.5 at 1526.