Confirm Action

Are you sure you want to proceed?

Model intelligence profile

Grok 4.5 benchmark results and model details

This family profile brings together 9 published benchmark records for xAI Grok 4.5. Every result keeps its original benchmark, configuration, scale, and source; unrelated scores are never averaged.

Published benchmark evidence available · Business Skills V3 not yet authorized

Provider

xAI

Base model family

Release date

2026-07-08

Exact-family source only

Published benchmarks

9

Native scales kept separate

Artificial Analysis

Reference available

As of 2026-07-16

Benchmark-level evidence

Published Grok 4.5 benchmark results

These are individual benchmarks, not collection rollups. Scores remain on their original scales, and agent or harness results stay labelled as system configurations.

How sources are reviewed →

Spring Prompt benchmark

BulletBench

3 configurations

60-second bullet Elo

Chess decision quality under a shared 60-second game clock and a 10+1 per-move clock.

Ladder Elo: anchored to a Stockfish skill ladder (random mover = 400). Internally consistent ordering, not FIDE-calibrated.

3 published configurations

Grok 4.5 (high)

0

Bullet 95% interval
0–84
10+1 Elo
0
Bullet games
96
Bullet median move time
4.13s
Average bullet game cost
$0.0390

Grok 4.5 (low)

0

Bullet 95% interval
0–118
10+1 Elo
0
Bullet games
96
Bullet median move time
4.37s
Average bullet game cost
$0.0368

Grok 4.5 (medium)

0

Bullet 95% interval
0–62
10+1 Elo
0
Bullet games
96
Bullet median move time
4.81s
Average bullet game cost
$0.0335
Source: Spring Prompt Recorded 3 Aug 2026 Ladder v1 · controls 10+1 and 60 Source record →

Spring Prompt benchmark

ROASBench

3 configurations

Average ROASBench score

A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.

3 published configurations

xAI: Grok 4.5

32.42

Average contribution profit
$204,735
Average ROAS
1.63×
Completed runs
3
Score variability
±1.31

xAI: Grok 4.5 · Max

32.44

Average contribution profit
$242,973
Average ROAS
1.67×
Completed runs
3
Score variability
±2.86

xAI: Grok 4.5 · Medium

32.95

Average contribution profit
$209,084
Average ROAS
1.63×
Completed runs
3
Score variability
±2.20
Source: Spring Prompt Recorded 25 Jul 2026 Published ROASBench cache · 12-month simulation Source record →

source-native benchmark

Artificial Analysis Intelligence Index

54

A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.

1 published configuration

Grok 4.5 (high)

54

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Intelligence Index v4.1 · 2026-07-30 Source record →

source-native benchmark

Coding Agent Configurations

57.9

Coding Agent Index v1.2 point score

Artificial Analysis Coding Agent Index v1.2 point scores and task-specific operational measurements for exact agent and model configurations.

These are exact agent-plus-model configurations, not bare-model scores.

1 published configuration

Grok Build - Grok 4.5 (high)

57.9

Mean cost per task
$2.59
Mean wall time
16.5 min
Source: Artificial Analysis Coding Agent Index Recorded 20 Jul 2026 V1.2 Source record →

source-native benchmark

GDPval-AA v2

51%

Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.

1 published configuration

Grok 4.5 (high)

51%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

GPQA Diamond

93%

The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.

1 published configuration

Grok 4.5 (high)

93%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

Humanity's Last Exam

40%

A broad expert-level benchmark of difficult academic reasoning and knowledge questions.

1 published configuration

Grok 4.5 (high)

40%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

MMMU-Pro

80%

A multimodal academic reasoning benchmark designed to reduce shortcuts and guessing across many disciplines.

1 published configuration

Grok 4.5 (high)

80%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

Terminal-Bench v2.1

82%

A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.

1 published configuration

Grok 4.5 (high)

82%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

Independent operational reference

Artificial Analysis details

Representative configuration: Grok 4.5 (high). Lifecycle: active; released 2026-07-08.

Input price

$2.00 / 1M tokens

Output price

$6.00 / 1M tokens

Median output speed

116.3 tok/s

Operational reference only. Price and throughput are not model-quality scores and never affect Spring Prompt comparisons.

Early-user signal

Initial community opinions

Anecdotal · never scored

A paraphrased editorial synthesis of 4 Reddit discussions created in the 14 days after launch (8 Jul 2026 to 22 Jul 2026). Anecdotal context only — it never affects scores or rankings.

The search-indexed launch-window sample for Grok 4.5 centered mainly on Cursor and Grok Build rather than a clearly identifiable consumer-chat rollout. Users often liked its price and speed but disagreed about sustained coding reliability and which host exposed the intended behavior. Because model identity was sometimes ambiguous, app-based quality changes should be treated cautiously.

Independent early-test signal

Early technical field tests

X · anecdotal · never scored

An editorial paraphrase of 2 launch-window field tests from 2 independent authors on X, including 2 reports with a described method or inspectable artifact. The window runs from 8 Jul 2026 up to 22 Jul 2026, and exact model identity was editor reviewed.

Early Grok 4.5 reports showed useful but uneven visual and asset work. One tester used it for a Three.js apartment scene and liked the overall feel while flagging errors in object positions and orientations. Another used it for asset sourcing, management, quality control, and fixes in a developing space game. These artifacts suggest practical creative support rather than dependable end-to-end scene construction.

These reports are selectively surfaced and are not a representative sample. They never affect benchmark scores, rankings, winners, or comparison outcomes.

Business Skills V3 · proposed

How Spring Prompt plans to test Grok 4.5

The setup below is proposed and may change until preflight and execution approval are complete. Existing benchmark evidence above does not authorize or stand in for a Business Skills V3 result.

Proposed reasoning
Provider-default reasoning
Provider revision
Exact provider revision will be resolved and frozen only after preflight and execution approval
Tools and service
No tools · Provider-default service tier
Planning configuration
grok-4.5
Generation controls
Temperature 0 proposed

Future first-party coverage

15 proposed Business Skills V3 task areas

These links describe evaluation contracts, not published Grok 4.5 results.

Model comparisons

Compare Grok 4.5 side by side

Editorially reviewed comparisons appear first. Every page matches only benchmark records with the same reviewed protocol key and does not manufacture an overall winner.

Publication safeguards

What must pass before a V3 result appears

  1. 1First-party task-local comparisons and eligible external evidence must both be present.
  2. 2Model and provider configuration identity must match the reviewed release exactly.
  3. 3Coverage, reliability, judge diagnostics, and sealed stability checks must pass.
  4. 4Uncertainty and missing evidence remain visible when results are published.

Stable family URL

Evidence can grow without changing the page

New reviewed benchmark snapshots, operational facts, community themes, and early field-test syntheses can be added here while the canonical model-family identity remains fixed.

More from xAI

Other owned model profiles