Grok 4.5 is an AI model released by SpaceXAI on Wednesday, Jul 8 2026, 82 days after Grok 4.3 Beta. Benchmark results (shown below) cover BullshitBench v2, SWE-Bench Pro, SWE-Bench Multilingual, DeepSWE 1.0, Next.js Evals, Terminal-Bench 2.1, and 3 more.
Benchmarks
Nonsense detection
BullshitBench v2Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
54%+4%vs Grok 4.3 Beta
Agentic coding
SWE-Bench ProCan the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better.
64.7%
Multilingual coding
SWE-Bench MultilingualLike SWE-Bench, but the coding problems span many programming languages, not just one. Tests how broadly the AI can code. Higher is better.
78%
Agentic coding
DeepSWE 1.0Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. Higher is better.
62%
Next.js coding
Next.js EvalsVercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better.
83%
Agentic terminal coding
Terminal-Bench 2.1Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better.
83.3%
Agentic legal work
Harvey's Legal Agent BenchmarkHarvey's test of whether an AI agent can complete real legal work — drafting and reviewing documents, working with spreadsheets and presentations, and navigating files the way a lawyer's assistant would. Higher is better.
12.92%
Medical admin work
MedScribeCan the AI support doctors with their administrative work, like notes and paperwork? Created by Vals AI. Higher is better.
86.88%
Community preference (code)
Arena Elo (Code)Like the text arena, but people vote on which AI writes better code. The votes become a chess-style Elo rating on arena.ai. Higher is better.
1566
About Grok 4.5
Grok 4.5, released July 8, 2026, is xAI's current flagship and its strongest coding model to date: a WebDev Arena Elo of 1566, 64.7% on SWE-Bench Pro, 78.0% on SWE-Bench Multilingual, and 83.3% on Terminal-Bench 2.1 — second only to Claude Fable 5 on the terminal-agent benchmark at release.
It also posted 86.88% on MedScribe for clinical documentation and 62.0% on DeepSWE 1.0. Arriving almost exactly a year after Grok 4, it squares off against Claude Fable 5, GPT-5.6, and Gemini 3.5 in the mid-2026 frontier cohort.