GLM-5.3-FlashOpen Weight
Released
GLM-5.3-Flash is an AI model released by Z.ai on Wednesday, Aug 26 2026, 12 days after GLM-5.3. It is an open-weight model — the trained weights are available to download and run. It is a 320B parameter model with a 1M token context window. Benchmark results (shown below) cover Toolathlon-Verified, CharXiv Reasoning, OfficeQA Pro, MVBench, MMVU, DeepSWE 1.1, and 9 more.
Available from
| DeepInfra | $0.075 | $0.25 | 1M | fp4 |
|---|---|---|---|---|
| Relace | $0.09 | $0.30 | 1M | — |
| Morph | $0.10 | $0.35 | 1M | fp8 |
| Wafer | $0.10 | $0.35 | 1M | — |
| StreamLake | $0.1124 | $0.3745 | 1M | fp8 |
| GMICloud | $0.1125 | $0.375 | 1M | fp8 |
| Novita | $0.132 | $0.44 | 1M | fp8 |
| Makora | $0.14 | $0.47 | 1M | — |
| BaseTen | $0.15 | $0.50 | 1M | fp8 |
| Cloudflare | $0.15 | $0.50 | 1.3M | — |
| CoreWeave | $0.15 | $0.50 | 1M | fp8 |
| Crusoe | $0.15 | $0.50 | 1M | fp4 |
| DigitalOcean | $0.15 | $0.50 | 1M | — |
| Fireworks | $0.15 | $0.50 | 1M | — |
| Friendli | $0.15 | $0.50 | 1M | — |
| Io Net | $0.15 | $0.50 | 256K | fp8 |
| Parasail | $0.15 | $0.50 | 1M | fp8 |
| Phala | $0.15 | $0.50 | 1M | fp8 |
| Reka | $0.15 | $0.50 | 256K | fp8 |
| Sail Research | $0.15 | $0.50 | 1M | fp8 |
| SiliconFlow | $0.15 | $0.50 | 1M | fp8 |
| Together | $0.15 | $0.50 | 1M | — |
| Venice | $0.15 | $0.50 | 1M | — |
| NextBit | $0.177 | $0.59 | 1M | fp8 |
| Modal | $0.45 | $1.50 | 1M | fp8 |
Benchmarks
Coding
Terminal & CLI
Agentic & tool use
Reasoning & science
Knowledge work
Multimodal
Community preference
About
GLM-5.3-Flash, released August 26, 2026, had already been running on OpenRouter as an uncredited stealth model called "Ox Alpha" when Z.ai claimed it as a GLM release earlier that day. At 320B total parameters with 18B active it was a fraction of the size of the 743B GLM-5.3 that preceded it by twelve days, and Z.ai benchmarked it against GLM-5.2 rather than that model: it beat GLM-5.2 on all eight tests the older model reported, most heavily on AutomationBench at 48.8% and DeepSWE v1.1 at 63.4%.
Across the fourteen benchmarks in the launch table it topped four outright — 78.4% on Toolathlon Verified, a GDPval-AA v2 rating of 1773, 62.4% on OfficeQA Pro and 78.0% on Chartography — and sat second or third on most of the rest, a few points behind GPT-5.6 Terra on the coding rows, Claude Opus 4.8 on Humanity's Last Exam with tools and CharXiv, and Gemini 3.7 Flash on the video tests. The much larger GLM-5.3 still led it on three of the five benchmarks the two shared.
Two things made the release notable beyond the scores. It was natively multimodal with a 1M-token context window and went out under the MIT License with weights on Hugging Face, a pointed contrast at the time with GLM-5.3, which had launched through the GLM Coding Plan and ZCode with no weights at all and did not get them until two days after Flash shipped. And Z.ai said the model had run entirely on Chinese AI chips through its Ox Alpha preview — a claim about serving infrastructure rather than training, but a pointed one at this size, and a data point about how far domestic silicon had come. Access opened the same day across the API, ZCode, chat.z.ai and AutoClaw.
Compare GLM-5.3-Flash with
Suggested comparisons
Frequently asked questions
GLM-5.3-Flash was released by Z.ai on Wednesday, Aug 26 2026.