LLaMA 3.1vsGrok 4.6

LLaMA 3.1
Grok 4.6
Benchmarks
Nonsense detection
BullshitBench v2
14%
Agentic coding
CursorBench v3.2
69.9%
Agentic coding
DeepSWE 1.1
65.9%
Agentic coding
FrontierCode v1.1 (Extended) · extended split
61.3%
Expert software engineering
APEX-SWE
56.4%
Agentic terminal coding
Terminal-Bench 3.0
26%
Expert agentic work
APEX-Agents
57.5%
Agentic legal work
Harvey's Legal Agent Benchmark
15.8%
Overall intelligence
AA Intelligence Index
61
Knowledge work
GDPval-AA v2
1753
Knowledge work
AA-Briefcase
1577
Overview
CompanyMetaSpaceXAI
Release dateJul 23 2024Aug 12 2026
AccessOpen WeightProprietary

Which is better: LLaMA 3.1 or Grok 4.6?

LLaMA 3.1 and Grok 4.6 don't publish scores on any of the same benchmarks, so there's no direct head-to-head comparison. LLaMA 3.1 shipped 750 days before Grok 4.6, so benchmark comparisons should account for the intervening progress.

LLaMA 3.1 is open weight, while Grok 4.6 is proprietary.

Direct benchmark comparisons are unavailable — LLaMA 3.1 and Grok 4.6 don't publish scores on any of the same benchmarks.

Frequently asked questions

LLaMA 3.1 was released by Meta on Jul 23 2024.

Grok 4.6 was released by SpaceXAI on Aug 12 2026.

LLaMA 3.1 is an open weight model released by Meta. Grok 4.6 is a proprietary model released by SpaceXAI.

Other comparisons