GLM-5 is an AI model released by Z.ai on Thursday, Feb 12 2026, 52 days after GLM-4.7. It is an open-weight model — the trained weights are available to download and run. It is a 744B parameter model. Benchmark results (shown below) cover BullshitBench v2, SWE-Bench Verified, SWE-Bench Multilingual, Terminal-Bench 2.0, BrowseComp, Humanity's Last Exam, and 2 more.
Benchmarks
Nonsense detection
BullshitBench v2Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
28%
Coding
SWE-Bench VerifiedReal coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better.
77.8%
Multilingual coding
SWE-Bench MultilingualLike SWE-Bench, but the coding problems span many programming languages, not just one. Tests how broadly the AI can code. Higher is better.
73.3%
Agentic terminal coding
Terminal-Bench 2.0Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? (Version 2.0 of the test.) Higher is better.
56.2%
Web browsing
BrowseCompCan the AI browse the web and track down hard-to-find answers? Higher is better.
75.9%
Multidisciplinary reasoning
Humanity's Last ExamHumanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better.
50.4%
with tools
Science
GPQA DiamondGraduate-level science questions in biology, physics, and chemistry — hard enough that subject-matter PhDs score around 65%. Higher is better.
86%
Community preference (code)
Arena Elo (Code)Like the text arena, but people vote on which AI writes better code. The votes become a chess-style Elo rating on arena.ai. Higher is better.
1430
About GLM-5
GLM-5, released February 12, 2026, doubled Z.ai's flagship to 744B parameters while keeping the weights open. It scored 77.8% on SWE-Bench Verified, 86.0% on GPQA, and 75.9% on BrowseComp — at release the strongest agentic-research score of any open-weight model — with 50.4% on Humanity's Last Exam with tools.
Landing a week after Claude Opus 4.6 and GPT-5.3-Codex, GLM-5 kept open weights within striking distance of the closed frontier through early 2026. GLM-5.1 followed in April with a 200K context window, and GLM-5.2 pushed the family to a 1M-token context in June 2026.