Confirm Action

Are you sure you want to proceed?

Independent model comparison

GPT-5.6 Luna vs Grok 4.5: evidence side by side

Compare the models using only evidence that can be identified and attributed. Exact benchmark matches appear first; independent facts and user reports remain separate and do not create an overall ranking.

Exact comparison coverage

7 benchmark records share an exact reviewed comparison key.

Evidence strength

Editorial review pending

Substantial matched evidence exists, but this pair has not passed the pair-native publication gate.

Models at a glance

Key differences

Detail GPT-5.6 Luna Grok 4.5
Provider

OpenAI

xAI

Release date Shown only where an exact reviewed release date is available.

Not verified

2026-07-08

Input price Artificial Analysis operational reference; see the attributed details below.

Not verified

$2.00 / 1M tokens

Output price Artificial Analysis operational reference; see the attributed details below.

Not verified

$6.00 / 1M tokens

Median output speed Artificial Analysis operational reference; see the attributed details below.

Not verified

116.3 tok/s

Exact benchmark coverage Only records carrying the same reviewed comparison identity are counted.

7 shared records

7 shared records

Price and speed appear only when an exact reviewed operational reference is attached to that model identity.

Primary comparison evidence

Exact matched benchmark evidence

7 exact matches

A shared benchmark name is not enough. A result appears side by side only when both model records carry the same reviewed comparison key. Duplicate or unversioned records are omitted rather than guessed into alignment.

Artificial Analysis Intelligence Index

A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.

GPT-5.6 Luna

46

Artificial Analysis Intelligence Index

Configuration: GPT-5.6 Luna (high)

Artificial Analysis · as of 30 Jul 2026 · Intelligence Index v4.1 · 2026-07-30

Grok 4.5

54

Artificial Analysis Intelligence Index

Configuration: Grok 4.5 (high)

Artificial Analysis · as of 30 Jul 2026 · Intelligence Index v4.1 · 2026-07-30

GDPval-AA v2

Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.

GPT-5.6 Luna

48%

GDPval-AA v2

Configuration: GPT-5.6 Luna (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Grok 4.5

51%

GDPval-AA v2

Configuration: Grok 4.5 (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

GPQA Diamond

The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.

GPT-5.6 Luna

89%

GPQA Diamond

Configuration: GPT-5.6 Luna (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Grok 4.5

93%

GPQA Diamond

Configuration: Grok 4.5 (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Humanity's Last Exam

A broad expert-level benchmark of difficult academic reasoning and knowledge questions.

GPT-5.6 Luna

32%

Humanity's Last Exam

Configuration: GPT-5.6 Luna (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Grok 4.5

40%

Humanity's Last Exam

Configuration: Grok 4.5 (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

MMMU-Pro

A multimodal academic reasoning benchmark designed to reduce shortcuts and guessing across many disciplines.

GPT-5.6 Luna

78%

MMMU-Pro

Configuration: GPT-5.6 Luna (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Grok 4.5

80%

MMMU-Pro

Configuration: Grok 4.5 (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

ROASBench

A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.

GPT-5.6 Luna

4 configurations

Average ROASBench score

Configuration: 4 configurations compared separately

Spring Prompt · Reviewed public-catalogue SQLite projection · 12-month simulation

Grok 4.5

3 configurations

Average ROASBench score

Configuration: 3 configurations compared separately

Spring Prompt · Reviewed public-catalogue SQLite projection · 12-month simulation

Terminal-Bench v2.1

A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.

GPT-5.6 Luna

70%

Terminal-Bench v2.1

Configuration: GPT-5.6 Luna (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Grok 4.5

82%

Terminal-Bench v2.1

Configuration: Grok 4.5 (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Secondary context · not head-to-head

Independent signals for each model

The sources below evaluated or discussed each model independently. Putting them in adjacent columns makes them easier to inspect, but it does not make their metrics, samples, or observations directly comparable.

Attributed third-party facts

Artificial Analysis details

These facts retain their source, date, protocol, and model identity. They are contextual model-level signals—not a Spring Prompt overall score.

GPT-5.6 Luna

No reviewed Artificial Analysis record is attached to this model snapshot.

Grok 4.5

Input price

$2.00 / 1M tokens

Representative configuration on the cited source page.

Artificial Analysis · as of 16 Jul 2026

Output price

$6.00 / 1M tokens

Representative configuration on the cited source page.

Artificial Analysis · as of 16 Jul 2026

Median output speed

116.3 tok/s

Median output tokens received per second after generation begins; this excludes time to first token and is not end-to-end response latency.

Artificial Analysis · as of 16 Jul 2026

Anecdotal model-level reports

Initial community opinions from Reddit

Recurring themes from each model's declared observation window. They are not a representative survey and are not direct comparisons between these models.

GPT-5.6 Luna

No editor-reviewed Reddit opinions paragraph is active for this model snapshot.

Grok 4.5

Initial community opinions

The search-indexed launch-window sample for Grok 4.5 centered mainly on Cursor and Grok Build rather than a clearly identifiable consumer-chat rollout. Users often liked its price and speed but disagreed about sustained coding reliability and which host exposed the intended behavior. Because model identity was sometimes ambiguous, app-based quality changes should be treated cautiously.

Source venue: Reddit · window: 2026-07-08 to 2026-07-22 · observed items: 4

Independent early-test reports

Early technical field tests from X

Editor-paraphrased reports from independent authors in each model's fixed 14-day launch window. These selectively surfaced anecdotes are not direct head-to-head tests unless they also appear in the dedicated direct-report section above.

GPT-5.6 Luna

No editor-reviewed X field-test paragraph is active for this model snapshot.

Grok 4.5

Early technical field tests

Early Grok 4.5 reports showed useful but uneven visual and asset work. One tester used it for a Three.js apartment scene and liked the overall feel while flagging errors in object positions and orientations. Another used it for asset sourcing, management, quality control, and fixes in a developing space game. These artifacts suggest practical creative support rather than dependable end-to-end scene construction.

Source venue: X · window: 2026-07-08 to 2026-07-22 · 2 tests · 2 independent authors · 2 with a method or artifact · identity reviewed

X field tests remain separate from Reddit community opinions, Artificial Analysis facts, and every scored benchmark.

How to read this comparison

Spring Prompt does not name an overall winner from unrelated benchmark scores, third-party metrics, Reddit opinions, or X field-test reports.

Different benchmarks measure different constructs and may use different populations, prompts, harnesses, model revisions, tools, and score scales. This page does not average those values or infer a winner from source coverage. Test both models on your own production work before making a consequential choice.