GLM-5.3-FlashOpen Weight
Released
GLM-5.3-Flash is an AI model released by Z.ai on Wednesday, Aug 26 2026, 12 days after GLM-5.3. It is an open-weight model — the trained weights are available to download and run. It is a 320B parameter model with a 1M token context window. Benchmark results (shown below) cover Toolathlon-Verified, CharXiv Reasoning, OfficeQA Pro, MVBench, MMVU, CursorBench 4.0, and 10 more.
Get Z.ai releases and AI news.
Every other Monday
Available from
| StreamLake | — | — | $0.069 | $0.23 | 1M | fp8 |
|---|---|---|---|---|---|---|
| DeepInfra | — | — | $0.075 | $0.25 | 1M | fp4 |
| Novita | — | — | $0.084 | $0.28 | 1M | fp8 |
| GMICloud | — | — | $0.09 | $0.30 | 1M | fp8 |
| Decart | — | — | $0.093 | $0.31 | 1M | fp4 |
| Near AI | — | — | $0.105 | $0.35 | 1M | fp8 |
| Phala | — | — | $0.1125 | $0.375 | 1M | nvfp4 |
| OpenInference | — | — | $0.044 | $0.45 | 1M | fp4 |
| Relace | — | — | $0.04 | $0.50 | 1M | — |
| InferenceNet | — | — | $0.07 | $0.50 | 1M | fp4 |
| Wafer | — | — | $0.07 | $0.50 | 1M | — |
| BaseTen | — | — | $0.15 | $0.50 | 1M | fp8 |
| CoreWeave | — | — | $0.15 | $0.50 | 1M | nvfp4 |
| Crusoe | — | — | $0.15 | $0.50 | 1M | fp4 |
| DigitalOcean | — | — | $0.15 | $0.50 | 1M | — |
| Fireworks | — | — | $0.15 | $0.50 | 1M | — |
| Friendli | — | — | $0.15 | $0.50 | 1M | — |
| Modal | — | — | $0.15 | $0.50 | 1M | nvfp4 |
| Parasail | — | — | $0.15 | $0.50 | 1M | fp4 |
| Together | — | — | $0.15 | $0.50 | 1M | — |
| Venice | — | — | $0.15 | $0.50 | 1M | — |
| Z.AI | — | — | $0.15 | $0.50 | 1M | fp8 |
| Inceptron | — | — | $0.12 | $0.55 | 1M | fp8 |
| Sail Research | — | — | $0.045 | $0.60 | 1M | fp4 |
| Sail Research | us | — | $0.045 | $0.60 | 1M | fp4 |
| Parasail | — | fast | $0.1875 | $0.625 | 1M | fp4 |
| Morph | — | — | $0.117 | $0.75 | 1M | fp8 |
| Fireworks | us | — | $0.225 | $0.75 | 1M | — |
| DekaLLM | — | — | $0.10 | $1.00 | 1M | — |
| Reka | — | — | $0.06 | $1.60 | 256K | — |
Benchmarks
Coding
Terminal & CLI
Agentic & tool use
Reasoning & science
Knowledge work
Multimodal
Community preference
Source: CursorBench, retrieved 8 October 2026. Other sources are identified on the linked benchmark pages.
Compare GLM-5.3-Flash with
Suggested comparisons
GLM-5.3-FlashvsGLM-5.3GLM-5.3-FlashvsGPT-6.1 SolGLM-5.3-FlashvsClaude Haiku 5.5GLM-5.3-FlashvsGemini 4 ArgonGLM-5.3-FlashvsMuse Spark 1.3GLM-5.3-FlashvsGrok 4.7GLM-5.3-FlashvsDeepSeek-V4.1-FlashGLM-5.3-FlashvsMistral Large 4GLM-5.3-FlashvsKimi K3GLM-5.3-FlashvsQwen3.8-Max-0902GLM-5.3-FlashvsNemotron 3.5 LightningAbout
GLM-5.3-Flash, released August 26, 2026, had already been running on OpenRouter as an uncredited stealth model called "Ox Alpha" when Z.ai claimed it as a GLM release earlier that day. At 320B total parameters with 18B active it was a fraction of the size of the 743B GLM-5.3 that preceded it by twelve days, and Z.ai benchmarked it against GLM-5.2 rather than that model: it beat GLM-5.2 on all eight tests the older model reported, most heavily on AutomationBench at 48.8% and DeepSWE v1.1 at 63.4%.
Across the fourteen benchmarks in the launch table it topped four outright — 78.4% on Toolathlon Verified, a GDPval-AA v2 rating of 1773, 62.4% on OfficeQA Pro and 78.0% on Chartography — and sat second or third on most of the rest, a few points behind GPT-5.6 Terra on the coding rows, Claude Opus 4.8 on Humanity's Last Exam with tools and CharXiv, and Gemini 3.7 Flash on the video tests. The much larger GLM-5.3 still led it on three of the five benchmarks the two shared.
Two things made the release notable beyond the scores. It was natively multimodal with a 1M-token context window and went out under the MIT License with weights on Hugging Face, a pointed contrast at the time with GLM-5.3, which had launched through the GLM Coding Plan and ZCode with no weights at all and did not get them until two days after Flash shipped. And Z.ai said the model had run entirely on Chinese AI chips through its Ox Alpha preview — a claim about serving infrastructure rather than training, but a pointed one at this size, and a data point about how far domestic silicon had come. Access opened the same day across the API, ZCode, chat.z.ai and AutoClaw.
Frequently asked questions
GLM-5.3-Flash was released by Z.ai on Wednesday, Aug 26 2026.