SpaceXAI's (formerly xAI) earlier 2026 flagship. Now superseded by Grok 4.6, the Opus-class model that holds the $2/$6 flagship rate as of August 12, 2026.
Grok 4.20 has been superseded by Grok 4.6, SpaceXAI's (formerly xAI) current Opus-class flagship, at the same $2/$6 rate. This page is kept for reference.
Grok 4.20 is xAI's 2026 flagship and its bid to compete on price as well as personality. Where the legacy Grok 4 launched at $3/$15 — premium-tier pricing, and it remains available at that rate — 4.20 lands at $2 input and $6 output per 1M tokens, undercutting every other flagship on our API table. It's competitive with frontier models on benchmarks, and the X (Twitter) integration remains the genuinely unique feature: real-time access to the timeline that no other provider has.
The model voice is also distinct — less filtered, more conversational — which is a feature or a bug depending on your product. For consumer apps that want personality, it's a draw; for enterprise document work, the Claude and GPT families remain the safer defaults.
Two lighter models share the family, both since repriced with tiered rates. Grok 4.3 ($1.25/$2.50 per 1M under 200K tokens, $2.50/$5 at 200K and above) is xAI's reasoning workhorse and the direct successor to the now-retired Grok 4, with a 1M token context window. Grok Build 0.1 ($1/$2 per 1M under 200K, $2/$4 at 200K and above) targets code tasks and replaced the retired Grok Code Fast 1 at the same standard-tier price point. Both trade flagship-grade reasoning for a lower price — use 4.20 or 4.5 when the answer needs to be smart, the lighter siblings when it needs to be cheap.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Grok 4.20 | $2 | $6 | 2M |
| Grok 4.3 (<200K) | $1.25 | $2.50 | 1M |
| Grok Build 0.1 (<200K) | $1 | $2 | 256K |
| Gemini 3.1 Pro | $2 | $12 | 1M |
| Claude Sonnet 5 | $2 | $10 | 1M |
At $2/$6, Grok 4.20 matches Gemini 3.1 Pro on input and halves it on output, while Sonnet 5 costs 2.5× as much for output. The honest trade: those rivals have stronger track records on careful long-form work and bigger ecosystems. Grok wins on price, real-time data, and voice; it has the most to prove on reliability-critical workloads.