# GLM-5.3

GLM-5.3 is an AI model released by Z.ai on Aug 14 2026. It has 743B parameters. At release it scored 28.3% on Terminal-Bench 3.0, 54.4% on ExploitBench and 130 on ExploitGym (6-hour budget).

## Facts

| Field | Value |
| --- | --- |
| Model | GLM-5.3 |
| Developer | Z.ai |
| Release date | Friday, Aug 14 2026 |
| Licensing | Proprietary |
| Parameters | 743B |

## Benchmark scores published at release

| Benchmark | Score | What it measures |
| --- | --- | --- |
| DeepSWE 1.1 | 66.9% | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. |
| Terminal-Bench 3.0 | 28.3% | command-line task completion (v3.0, much harder task set) |
| CyberGym | 84.5% | Tests the AI on cybersecurity challenges — finding and exploiting software weaknesses inside a safe sandbox. Higher is better. |
| ExploitBench | 54.4% | A 'capability ladder' for security research, built by CMU researchers: the AI is given known bugs in Chrome's V8 engine and scored on how far it gets toward a working exploit inside a research sandbox — from understanding the patch to triggering a crash. Higher is better. |
| ExploitGym (6-hour budget) | 130 | Can an AI agent turn a known software vulnerability into a working attack in a controlled lab? Built by MPI-SP researchers, the score is how many of 898 real cases (userspace programs, the V8 engine, the Linux kernel) it cracks — here with a 6-hour compute budget per case. Higher is better. |
| ExploitGym (2-hour budget) | 105 | Can an AI agent turn a known software vulnerability into a working attack in a controlled lab? Built by MPI-SP researchers, the score is how many of 898 real cases (userspace programs, the V8 engine, the Linux kernel) it cracks — here with a 2-hour compute budget per case. Higher is better. |
| Humanity's Last Exam (with tools) | 62.5% | Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better. |
| Agent's Last Exam | 28.5% | A hard set of desktop and operating-system tasks an AI agent has to finish by looking at the screen and working the machine itself. The score is the share it passes outright — partial credit does not count. Higher is better. |
| AutomationBench | 48.2% | Tests whether the AI can run real multi-step business workflows — the kind of end-to-end office processes companies want to automate — from start to finish. Higher is better. |
| GDPval-AA v2 | 1769 | economically valuable knowledge work (v2, re-based Elo) |

## About GLM-5.3

GLM-5.3, released August 14, 2026, was Z.ai's coding-and-cybersecurity push — "Built to Code. Ready for Cyber Defense." — built by post-training the 743B GLM-5-generation base rather than scaling it up. At release it scored 66.9% on DeepSWE, more than twenty points above GLM-5.2, with 62.5% on Humanity's Last Exam with tools and a GDPval-AA v2 rating of 1769 — at the time the best published figure from any open-model lab, and ahead of the Claude and GPT flagships in Z.ai's own launch comparison.

The cybersecurity positioning was the real novelty: 84.5% on CyberGym topped Z.ai's launch chart outright, and 54.4% on ExploitBench — CMU's exploit-development ladder — set a launch-time high among open-model labs, though the closed Anthropic and OpenAI flagships still led that test by over twenty points. Unusually for the GLM line, launch day came without weights: access started through the GLM Coding Plan and ZCode, with API access and open weights slated to follow in stages after safety evaluations. It was Z.ai's seventh flagship release in thirteen months.

## Questions and answers

### When was GLM-5.3 released?

GLM-5.3 was released by Z.ai on Friday, Aug 14 2026.

### Who made GLM-5.3?

GLM-5.3 was built by Z.ai. Chinese AI lab spun out of Tsinghua University (formerly Zhipu AI), building the open-weight GLM family. Rebranded internationally as Z.ai in 2025.

### What benchmark scores did GLM-5.3 get?

GLM-5.3 reports 10 tracked benchmark scores — DeepSWE 1.1: 66.9%; Terminal-Bench 3.0: 28.3%; CyberGym: 84.5%; ExploitBench: 54.4%; ExploitGym (6-hour budget): 130; ExploitGym (2-hour budget): 105; Humanity's Last Exam (with tools): 62.5%; Agent's Last Exam: 28.5%; AutomationBench: 48.2%; GDPval-AA v2: 1769. Scores are the figures published at release by Z.ai. It holds the best score among all models tracked here on Terminal-Bench 3.0, ExploitBench, ExploitGym (6-hour budget), ExploitGym (2-hour budget), Agent's Last Exam and AutomationBench.

### How many parameters does GLM-5.3 have?

GLM-5.3 is reported at 743B parameters.

### Is GLM-5.3 open source?

No. GLM-5.3 is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after GLM-5.3?

Z.ai's previous tracked release was GLM-5.2 on Jun 16 2026, 59 days earlier. It is the most recent Z.ai model tracked on AI Release Tracker.


---

Canonical page: https://aireleasetracker.com/model/zai/glm-5.3
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
