Provider
xAI
Base model family
Model intelligence profile
This family profile brings together 9 published benchmark records for xAI Grok 4.5. Every result keeps its original benchmark, configuration, scale, and source; unrelated scores are never averaged.
Provider
xAI
Base model family
Release date
2026-07-08
Exact-family source only
Published benchmarks
9
Native scales kept separate
Artificial Analysis
Reference available
As of 2026-07-16
Benchmark-level evidence
These are individual benchmarks, not collection rollups. Scores remain on their original scales, and agent or harness results stay labelled as system configurations.
Spring Prompt benchmark
3 configurations
60-second bullet Elo
Chess decision quality under a shared 60-second game clock and a 10+1 per-move clock.
Ladder Elo: anchored to a Stockfish skill ladder (random mover = 400). Internally consistent ordering, not FIDE-calibrated.
Grok 4.5 (high)
0
Grok 4.5 (low)
0
Grok 4.5 (medium)
0
Spring Prompt benchmark
3 configurations
Average ROASBench score
A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.
xAI: Grok 4.5
32.42
xAI: Grok 4.5 · Max
32.44
xAI: Grok 4.5 · Medium
32.95
source-native benchmark
54
A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.
Grok 4.5 (high)
54
source-native benchmark
57.9
Coding Agent Index v1.2 point score
Artificial Analysis Coding Agent Index v1.2 point scores and task-specific operational measurements for exact agent and model configurations.
These are exact agent-plus-model configurations, not bare-model scores.
Grok Build - Grok 4.5 (high)
57.9
source-native benchmark
51%
Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.
Grok 4.5 (high)
51%
source-native benchmark
93%
The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.
Grok 4.5 (high)
93%
source-native benchmark
40%
A broad expert-level benchmark of difficult academic reasoning and knowledge questions.
Grok 4.5 (high)
40%
source-native benchmark
80%
A multimodal academic reasoning benchmark designed to reduce shortcuts and guessing across many disciplines.
Grok 4.5 (high)
80%
source-native benchmark
82%
A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.
Grok 4.5 (high)
82%
Independent operational reference
Representative configuration: Grok 4.5 (high). Lifecycle: active; released 2026-07-08.
Input price
$2.00 / 1M tokens
Output price
$6.00 / 1M tokens
Median output speed
116.3 tok/s
Operational reference only. Price and throughput are not model-quality scores and never affect Spring Prompt comparisons.
Early-user signal
A paraphrased editorial synthesis of 4 Reddit discussions created in the 14 days after launch (8 Jul 2026 to 22 Jul 2026). Anecdotal context only — it never affects scores or rankings.
The search-indexed launch-window sample for Grok 4.5 centered mainly on Cursor and Grok Build rather than a clearly identifiable consumer-chat rollout. Users often liked its price and speed but disagreed about sustained coding reliability and which host exposed the intended behavior. Because model identity was sometimes ambiguous, app-based quality changes should be treated cautiously.
Independent early-test signal
An editorial paraphrase of 2 launch-window field tests from 2 independent authors on X, including 2 reports with a described method or inspectable artifact. The window runs from 8 Jul 2026 up to 22 Jul 2026, and exact model identity was editor reviewed.
Early Grok 4.5 reports showed useful but uneven visual and asset work. One tester used it for a Three.js apartment scene and liked the overall feel while flagging errors in object positions and orientations. Another used it for asset sourcing, management, quality control, and fixes in a developing space game. These artifacts suggest practical creative support rather than dependable end-to-end scene construction.
These reports are selectively surfaced and are not a representative sample. They never affect benchmark scores, rankings, winners, or comparison outcomes.
Business Skills V3 · proposed
The setup below is proposed and may change until preflight and execution approval are complete. Existing benchmark evidence above does not authorize or stand in for a Business Skills V3 result.
Future first-party coverage
These links describe evaluation contracts, not published Grok 4.5 results.
Model comparisons
Editorially reviewed comparisons appear first. Every page matches only benchmark records with the same reviewed protocol key and does not manufacture an overall winner.
Publication safeguards
Stable family URL
New reviewed benchmark snapshots, operational facts, community themes, and early field-test syntheses can be added here while the canonical model-family identity remains fixed.
More from xAI