GLM-5.3-FlashOpen Weight

Released

GLM-5.3-Flash is an AI model released by Z.ai on Wednesday, Aug 26 2026, 12 days after GLM-5.3. It is an open-weight model — the trained weights are available to download and run. It is a 320B parameter model with a 1M token context window. Benchmark results (shown below) cover Toolathlon-Verified, CharXiv Reasoning, OfficeQA Pro, MVBench, MMVU, CursorBench 4.0, and 10 more.

Available from

StreamLake——$0.069$0.231Mfp8
DeepInfra——$0.075$0.251Mfp4
Novita——$0.084$0.281Mfp8
GMICloud——$0.09$0.301Mfp8
Decart——$0.093$0.311Mfp4
Near AI——$0.105$0.351Mfp8
Phala——$0.1125$0.3751Mnvfp4
OpenInference——$0.044$0.451Mfp4
Relace——$0.04$0.501M—
InferenceNet——$0.07$0.501Mfp4
Wafer——$0.07$0.501M—
BaseTen——$0.15$0.501Mfp8
CoreWeave——$0.15$0.501Mnvfp4
Crusoe——$0.15$0.501Mfp4
DigitalOcean——$0.15$0.501M—
Fireworks——$0.15$0.501M—
Friendli——$0.15$0.501M—
Modal——$0.15$0.501Mnvfp4
Parasail——$0.15$0.501Mfp4
Together——$0.15$0.501M—
Venice——$0.15$0.501M—
Z.AI——$0.15$0.501Mfp8
Inceptron——$0.12$0.551Mfp8
Sail Research——$0.045$0.601Mfp4
Sail Researchus—$0.045$0.601Mfp4
Parasail—fast$0.1875$0.6251Mfp4
Morph——$0.117$0.751Mfp8
Fireworksus—$0.225$0.751M—
DekaLLM——$0.10$1.001M—
Reka——$0.06$1.60256K—
USD per 1M tokens

Benchmarks

Coding

CursorBench 4.0via CursorBench
36.8%
#13 of 16
DeepSWE 1.1
63.4%
#18 of 32
NL2Repo-Bench
56.3%
#3 of 6

Terminal & CLI

Terminal-Bench 2.1
84.3%
#12 of 32

Agentic & tool use

Toolathlon-Verified
78.4%
#1
Agent's Last Exam
26.3%pass@1
#6 of 9
AutomationBench
48.8%
#6 of 16

Reasoning & science

Humanity's Last Exam
55.3%with tools
#17 of 46

Knowledge work

GDPval-AA v2
1773
#3 of 20

Multimodal

CharXiv Reasoning
89.4%with tools
OfficeQA Pro
62.4%
MVBench
77.8%
MMVU
80.5%
Chartography
78%with tools
#5 of 5
BabyVision
53.4%
#3 of 4

Community preference

threejseval
1378
#18 of 27

Source: CursorBench, retrieved 8 October 2026. Other sources are identified on the linked benchmark pages.

Compare GLM-5.3-Flash with

GLM-5.3-Flash

About

GLM-5.3-Flash, released August 26, 2026, had already been running on OpenRouter as an uncredited stealth model called "Ox Alpha" when Z.ai claimed it as a GLM release earlier that day. At 320B total parameters with 18B active it was a fraction of the size of the 743B GLM-5.3 that preceded it by twelve days, and Z.ai benchmarked it against GLM-5.2 rather than that model: it beat GLM-5.2 on all eight tests the older model reported, most heavily on AutomationBench at 48.8% and DeepSWE v1.1 at 63.4%.

Across the fourteen benchmarks in the launch table it topped four outright — 78.4% on Toolathlon Verified, a GDPval-AA v2 rating of 1773, 62.4% on OfficeQA Pro and 78.0% on Chartography — and sat second or third on most of the rest, a few points behind GPT-5.6 Terra on the coding rows, Claude Opus 4.8 on Humanity's Last Exam with tools and CharXiv, and Gemini 3.7 Flash on the video tests. The much larger GLM-5.3 still led it on three of the five benchmarks the two shared.

Two things made the release notable beyond the scores. It was natively multimodal with a 1M-token context window and went out under the MIT License with weights on Hugging Face, a pointed contrast at the time with GLM-5.3, which had launched through the GLM Coding Plan and ZCode with no weights at all and did not get them until two days after Flash shipped. And Z.ai said the model had run entirely on Chinese AI chips through its Ox Alpha preview — a claim about serving infrastructure rather than training, but a pointed one at this size, and a data point about how far domestic silicon had come. Access opened the same day across the API, ZCode, chat.z.ai and AutoClaw.

Frequently asked questions

GLM-5.3-Flash was released by Z.ai on Wednesday, Aug 26 2026.

All Z.ai releases

13 tracked