LLaMA 3.1vsGrok‑2
LLaMA 3.1 | Grok‑2 | |
|---|---|---|
| API pricingUSD per 1M tokens · lower wins | ||
Cheapest inputLowest input rate across third-party providers, excluding the lab itself. The cheapest endpoint may run a quantised build or a shorter context — see "Available from" on the model page. | $0.02DeepInfra | — |
Cheapest outputLowest output rate across third-party providers, excluding the lab itself. May come from a different provider than the cheapest input. | $0.04DeepInfra | — |
| BenchmarksPublished by one model only | ||
BullshitBench v2Nonsense detection — Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better. | 14% | — |
| Overview | ||
| Company | Meta | SpaceXAI |
| Release date | Jul 23 2024 | Aug 14 2024 |
| Access | Open Weight | Proprietary |
Other comparisons
Frequently asked questions
LLaMA 3.1 and Grok‑2 don't publish scores on any of the same benchmarks, so there's no direct head-to-head comparison. LLaMA 3.1 shipped 22 days before Grok‑2, so benchmark comparisons should account for the intervening progress.
LLaMA 3.1 is open weight, while Grok‑2 is proprietary.
Direct benchmark comparisons are unavailable — LLaMA 3.1 and Grok‑2 don't publish scores on any of the same benchmarks.