# GLM-5.3-Flash

GLM-5.3-Flash is an AI model released by Z.ai on Aug 26 2026. It has 320B parameters, a 1M context window and open weight. Tracked results include 78.4% on Toolathlon-Verified, 89.4% on CharXiv Reasoning (with tools) and 62.4% on OfficeQA Pro.

## Facts

| Field | Value |
| --- | --- |
| Model | GLM-5.3-Flash |
| Developer | Z.ai |
| Release date | Wednesday, Aug 26 2026 |
| Licensing | Open Weight |
| Parameters | 320B |
| Context window | 1M |

## Tracked benchmark scores

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| CursorBench 4.0 | 36.8% | [CursorBench](https://cursor.com/cursorbench), retrieved 2026-10-08 | Cursor's own test of coding agents on ambiguous, multi-file tasks taken from real Cursor sessions — editing, refactoring, investigating a codebase, understanding what the user meant, managing jobs and following a design. Cursor runs each model at several reasoning efforts; each release here carries the score of its best listed effort. Scores aren't comparable with earlier CursorBench versions. Higher is better. |
| DeepSWE 1.1 | 63.4% | Lab | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. |
| NL2Repo-Bench | 56.3% | Lab | Tests whether the AI can turn a natural-language requirement into working code across an entire repository, not just produce a single function or patch. Higher is better. |
| Terminal-Bench 2.1 | 84.3% | Lab | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better. |
| Toolathlon-Verified | 78.4% | Lab | Tests how well the AI uses everyday personal tools and apps to get things done — a human-checked version of Toolathlon. Higher is better. |
| Humanity's Last Exam (with tools) | 55.3% | Lab | Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better. |
| Agent's Last Exam (pass@1) | 26.3% | Lab | A hard set of desktop and operating-system tasks an AI agent has to finish by looking at the screen and working the machine itself. The score is the share it passes outright — partial credit does not count. Higher is better. |
| AutomationBench | 48.8% | Lab | Tests whether the AI can run real multi-step business workflows — the kind of end-to-end office processes companies want to automate — from start to finish. Higher is better. |
| GDPval-AA v2 | 1773 | Lab | economically valuable knowledge work (v2, re-based Elo) |
| CharXiv Reasoning (with tools) | 89.4% | Lab | The same chart-and-figure reasoning test, run with the AI allowed to use tools — writing code to inspect the image, for instance — rather than reading the chart unaided. Scores run higher than the unaided version, so read the two as separate tests. Higher is better. |
| Chartography (with tools) | 78% | Lab | A chart-centred test run with tools available to the AI, reported separately from the chart-reading benchmarks above it. Higher is better. |
| OfficeQA Pro | 62.4% | Lab | Questions about office documents, where answering depends on reading the page as a document — layout, tables and figures included — rather than as loose text. Higher is better. |
| MVBench | 77.8% | Lab | Video questions that cannot be answered from any single frame: the AI has to follow what changes over time — the order things happen in, what moved where. Higher is better. |
| MMVU | 80.5% | Lab | Expert-level video questions drawn from specific disciplines, where answering means applying subject knowledge to what is happening on screen rather than just describing it. Higher is better. |
| BabyVision | 53.4% | Lab | Tests core visual reasoning — seeing and understanding images the way even young children can, which AIs often find surprisingly hard. Higher is better. |
| threejseval | 1378 | [threejseval](https://threejseval.com/ranking) | Every model gets the same prompt — "the Eiffel Tower", "a glass fishbowl", "a robot arm picking toys into a box" — and builds a 3D scene in Three.js. Real people then see two scenes side by side, names hidden, and vote for the one they prefer. The votes become a chess-style Elo rating on threejseval.com, averaged across all the prompts. It measures whether the scene looks and moves right to a human eye, not whether the code passes a test. Higher is better. |

## About GLM-5.3-Flash

GLM-5.3-Flash, released August 26, 2026, had already been running on OpenRouter as an uncredited stealth model called "Ox Alpha" when Z.ai claimed it as a GLM release earlier that day. At 320B total parameters with 18B active it was a fraction of the size of the 743B GLM-5.3 that preceded it by twelve days, and Z.ai benchmarked it against GLM-5.2 rather than that model: it beat GLM-5.2 on all eight tests the older model reported, most heavily on AutomationBench at 48.8% and DeepSWE v1.1 at 63.4%.

Across the fourteen benchmarks in the launch table it topped four outright — 78.4% on Toolathlon Verified, a GDPval-AA v2 rating of 1773, 62.4% on OfficeQA Pro and 78.0% on Chartography — and sat second or third on most of the rest, a few points behind GPT-5.6 Terra on the coding rows, Claude Opus 4.8 on Humanity's Last Exam with tools and CharXiv, and Gemini 3.7 Flash on the video tests. The much larger GLM-5.3 still led it on three of the five benchmarks the two shared.

Two things made the release notable beyond the scores. It was natively multimodal with a 1M-token context window and went out under the MIT License with weights on Hugging Face, a pointed contrast at the time with GLM-5.3, which had launched through the GLM Coding Plan and ZCode with no weights at all and did not get them until two days after Flash shipped. And Z.ai said the model had run entirely on Chinese AI chips through its Ox Alpha preview — a claim about serving infrastructure rather than training, but a pointed one at this size, and a data point about how far domestic silicon had come. Access opened the same day across the API, ZCode, chat.z.ai and AutoClaw.

## Questions and answers

### When was GLM-5.3-Flash released?

GLM-5.3-Flash was released by Z.ai on Wednesday, Aug 26 2026.

### Who made GLM-5.3-Flash?

GLM-5.3-Flash was built by Z.ai. Chinese AI lab spun out of Tsinghua University (formerly Zhipu AI), building the open-weight GLM family. Rebranded internationally as Z.ai in 2025.

### What benchmark scores did GLM-5.3-Flash get?

GLM-5.3-Flash reports 16 tracked benchmark scores — CursorBench 4.0: 36.8%; DeepSWE 1.1: 63.4%; NL2Repo-Bench: 56.3%; Terminal-Bench 2.1: 84.3%; Toolathlon-Verified: 78.4%; Humanity's Last Exam (with tools): 55.3%; Agent's Last Exam (pass@1): 26.3%; AutomationBench: 48.8%; GDPval-AA v2: 1773; CharXiv Reasoning (with tools): 89.4%; Chartography (with tools): 78%; OfficeQA Pro: 62.4%; MVBench: 77.8%; MMVU: 80.5%; BabyVision: 53.4%; threejseval: 1378. Tracked scores may come from lab reports or independent benchmarks; source details accompany the benchmark data. It holds the best score among all models tracked here on Toolathlon-Verified, CharXiv Reasoning (with tools), OfficeQA Pro, MVBench and MMVU.

### What is the context window of GLM-5.3-Flash?

GLM-5.3-Flash has a context window of 1M. That is the maximum amount of input plus output the model can hold in a single request.

### How many parameters does GLM-5.3-Flash have?

GLM-5.3-Flash is reported at 320B parameters.

### Is GLM-5.3-Flash open source?

Partly. GLM-5.3-Flash is an open-weight model: the trained weights are free to download and run locally or on your own infrastructure, but the training data and code are not fully released and the license may restrict some commercial uses. It is not open source in the strict sense.

### What came before and after GLM-5.3-Flash?

Z.ai's previous tracked release was GLM-5.3 on Aug 14 2026, 12 days earlier. It is the most recent Z.ai model tracked on AI Release Tracker.


---

Canonical page: https://aireleasetracker.com/model/zai/glm-5.3-flash
Site index: https://aireleasetracker.com/llms.txt
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
