xAI
Grok 4.20
- Benchmark records
- 1
- Snapshot
- 2026-04-02
Independent model comparison
Compare the models using only evidence that can be identified and attributed. Exact benchmark matches appear first; independent facts and user reports remain separate and do not create an overall ranking.
Exact comparison coverage
No benchmark records currently share an exact reviewed comparison key.
Evidence strength
No exact benchmark match
Model-level evidence is kept separate until a compatible comparison record is available.
Models at a glance
xAI
xAI
| Detail | Grok 4.20 | Grok 4.5 |
|---|---|---|
| Provider |
xAI |
xAI |
| Release date Shown only where an exact reviewed release date is available. |
Not verified |
2026-07-08 |
| Input price Artificial Analysis operational reference; see the attributed details below. |
Not verified |
$2.00 / 1M tokens |
| Output price Artificial Analysis operational reference; see the attributed details below. |
Not verified |
$6.00 / 1M tokens |
| Median output speed Artificial Analysis operational reference; see the attributed details below. |
Not verified |
116.3 tok/s |
| Exact benchmark coverage Only records carrying the same reviewed comparison identity are counted. |
0 shared records |
0 shared records |
Price and speed appear only when an exact reviewed operational reference is attached to that model identity.
Primary comparison evidence
A shared benchmark name is not enough. A result appears side by side only when both model records carry the same reviewed comparison key. Duplicate or unversioned records are omitted rather than guessed into alignment.
Each model may still have useful individual evidence on its profile, but those records cannot be treated as a direct comparison until their benchmark and protocol identities align.
Secondary context · not head-to-head
The sources below evaluated or discussed each model independently. Putting them in adjacent columns makes them easier to inspect, but it does not make their metrics, samples, or observations directly comparable.
Attributed third-party facts
These facts retain their source, date, protocol, and model identity. They are contextual model-level signals—not a Spring Prompt overall score.
No reviewed Artificial Analysis record is attached to this model snapshot.
Input price
$2.00 / 1M tokens
Representative configuration on the cited source page.
Artificial Analysis · as of 16 Jul 2026
Output price
$6.00 / 1M tokens
Representative configuration on the cited source page.
Artificial Analysis · as of 16 Jul 2026
Median output speed
116.3 tok/s
Median output tokens received per second after generation begins; this excludes time to first token and is not end-to-end response latency.
Artificial Analysis · as of 16 Jul 2026
Anecdotal model-level reports
Recurring themes from each model's declared observation window. They are not a representative survey and are not direct comparisons between these models.
No editor-reviewed Reddit opinions paragraph is active for this model snapshot.
Initial community opinions
The search-indexed launch-window sample for Grok 4.5 centered mainly on Cursor and Grok Build rather than a clearly identifiable consumer-chat rollout. Users often liked its price and speed but disagreed about sustained coding reliability and which host exposed the intended behavior. Because model identity was sometimes ambiguous, app-based quality changes should be treated cautiously.
Source venue: Reddit · window: 2026-07-08 to 2026-07-22 · observed items: 4
Independent early-test reports
Editor-paraphrased reports from independent authors in each model's fixed 14-day launch window. These selectively surfaced anecdotes are not direct head-to-head tests unless they also appear in the dedicated direct-report section above.
Early technical field tests
Early Grok 4.20 evidence came from two broad comparative exercises: a thirty-seven-model test of implicit understanding in text and a blind twenty-four-model OpenClaw evaluation. Those artifacts provide more structure than general launch reactions and make harness effects visible. They remain informal, tester-designed evaluations, so the evidence should be read as directional rather than representative.
Source venue: X · window: 2026-03-31 to 2026-04-14 · 2 tests · 2 independent authors · 2 with a method or artifact · identity reviewed
Early technical field tests
Early Grok 4.5 reports showed useful but uneven visual and asset work. One tester used it for a Three.js apartment scene and liked the overall feel while flagging errors in object positions and orientations. Another used it for asset sourcing, management, quality control, and fixes in a developing space game. These artifacts suggest practical creative support rather than dependable end-to-end scene construction.
Source venue: X · window: 2026-07-08 to 2026-07-22 · 2 tests · 2 independent authors · 2 with a method or artifact · identity reviewed
X field tests remain separate from Reddit community opinions, Artificial Analysis facts, and every scored benchmark.
Spring Prompt does not name an overall winner from unrelated benchmark scores, third-party metrics, Reddit opinions, or X field-test reports.
Different benchmarks measure different constructs and may use different populations, prompts, harnesses, model revisions, tools, and score scales. This page does not average those values or infer a winner from source coverage. Test both models on your own production work before making a consequential choice.