LLaMA 3.1vsGrok 4.5

LLaMA 3.1
Grok 4.5
Benchmarks
Nonsense detection
BullshitBench v2
14%
54%Best
Prompt injection robustness
Gray Swan IPI · k = 1
13.4%
Prompt injection robustness
Gray Swan IPI · k = 10
54.2%
Prompt injection robustness
Gray Swan IPI · k = 15
60.8%
Agentic coding
SWE-Bench Pro
64.7%
Multilingual coding
SWE-Bench Multilingual
78%
Agentic coding
CursorBench v3.2
66.7%
Agentic coding
DeepSWE 1.1
54%
Agentic coding
DeepSWE 1.0
62%
Next.js coding
Next.js Evals
83%
Agentic computer work
Frontier-Bench v0.1
17.8%
Agentic terminal coding
Terminal-Bench 2.1
83.3%
Agentic legal work
Harvey's Legal Agent Benchmark
12.92%
Medical admin work
MedScribe
86.88%
Community preference (code)
Arena Elo (Code)
1549
Overview
CompanyMetaSpaceXAI
Release dateJul 23 2024Jul 8 2026
AccessOpen WeightProprietary

Which is better: LLaMA 3.1 or Grok 4.5?

Grok 4.5 leads LLaMA 3.1 on 1 of the 1 benchmark they both report (BullshitBench v2). LLaMA 3.1 shipped 715 days before Grok 4.5, so benchmark comparisons should account for the intervening progress.

LLaMA 3.1 is open weight, while Grok 4.5 is proprietary.

On BullshitBench v2, Grok 4.5 leads at 54% vs LLaMA 3.1 at 14%.

Frequently asked questions

LLaMA 3.1 was released by Meta on Jul 23 2024.

Grok 4.5 was released by SpaceXAI on Jul 8 2026.

LLaMA 3.1 is an open weight model released by Meta. Grok 4.5 is a proprietary model released by SpaceXAI.

Other comparisons