Z.ai
GLM-5.1Open Weight
GLM-5.1 is an AI model released by Z.ai on Tuesday, Apr 7 2026, 54 days after GLM-5. It is an open-weight model — the trained weights are available to download and run. It is a 744B parameter model with a 200k token context window. Benchmark results (shown below) cover BullshitBench v2, Next.js Evals, Terminal-Bench 2.0, BrowseComp, Humanity's Last Exam, GPQA Diamond, and 1 more.
Benchmarks
Nonsense detection
BullshitBench v2Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
22%
Next.js coding
Next.js EvalsVercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better.
75%
Agentic terminal coding
Terminal-Bench 2.0Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? (Version 2.0 of the test.) Higher is better.
63.5%
Web browsing
BrowseCompCan the AI browse the web and track down hard-to-find answers? Higher is better.
68%
Multidisciplinary reasoning
Humanity's Last ExamHumanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better.
52.3%
with tools
Science
GPQA DiamondGraduate-level science questions in biology, physics, and chemistry — hard enough that subject-matter PhDs score around 65%. Higher is better.
86.2%
Community preference (code)
Arena Elo (Code)Like the text arena, but people vote on which AI writes better code. The votes become a chess-style Elo rating on arena.ai. Higher is better.
1527