Grok 4.20 Beta
Released
Grok 4.20 Beta is an AI model released by SpaceXAI on Tuesday, Feb 17 2026, 8 days after Composer 1.5. Benchmark results (shown below) cover BullshitBench v2, SWE-Bench Verified, and ARC-AGI-2.
Get SpaceXAI releases and AI news.
One email, every other Monday.
API pricing
Under 200K input tokens
- InputWhat you pay for everything you send the model — your question, plus any documents or earlier conversation you include with it.
- $1.25
- Cached inputA reduced rate for text you send over and over. If every request starts with the same instructions or the same document, the provider keeps a copy ready and charges less to read it again.
- $0.20
- OutputWhat you pay for the text the model writes back. It is normally the dearer half: producing an answer costs more than reading one.
- $2.50
At least 200K input tokens
- InputWhat you pay for everything you send the model — your question, plus any documents or earlier conversation you include with it.
- $2.50
- Cached inputA reduced rate for text you send over and over. If every request starts with the same instructions or the same document, the provider keeps a copy ready and charges less to read it again.
- $0.40
- OutputWhat you pay for the text the model writes back. It is normally the dearer half: producing an answer costs more than reading one.
- $5.00
- Reasoning or thinking is supported.
USD per 1M tokensEvery price here is for one million tokens. A token is roughly three-quarters of a word, so a million tokens is about 750,000 words of text.Model mappingVerified August 18, 2026Read from docs.x.ai. Click the check to open the source.
Benchmarks
Coding
SWE-Bench VerifiedCoding — Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better.via BenchLM
76.7%
#25 of 57Best: Claude Opus 5 · 96%
Reasoning & science
Robustness
BullshitBench v2Nonsense detection — Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
56%
#22 of 77Best: Claude Opus 4.8 · 95%
Source: BenchLM, retrieved 31 August 2026. Every other score here is the figure the lab published at launch.
Compare Grok 4.20 Beta with
Suggested comparisons
Grok 4.20 BetavsGrok 4.6Grok 4.20 BetavsGPT-6 AstraGrok 4.20 BetavsClaude Fable 5.1Grok 4.20 BetavsGemini 3.8 FlashGrok 4.20 BetavsMuse Spark 1.3Grok 4.20 BetavsDeepSeek-V4.1-FlashGrok 4.20 BetavsMistral Medium 3.5Grok 4.20 BetavsKimi K3Grok 4.20 BetavsGLM-5.3-FlashGrok 4.20 BetavsQwen3.8-Max-0902Grok 4.20 BetavsNemotron 3.5 LightningFrequently asked questions
Grok 4.20 Beta was released by SpaceXAI on Tuesday, Feb 17 2026.