Confirm Action

Are you sure you want to proceed?

Model intelligence profile

Claude Opus 4.8 benchmark results and model details

This family profile brings together 5 published benchmark records for Anthropic Claude Opus 4.8. Every result keeps its original benchmark, configuration, scale, and source; unrelated scores are never averaged.

Published benchmark evidence available · Business Skills V3 not yet authorized

Provider

Anthropic

Base model family

Release date

2026-05-28

Exact-family source only

Published benchmarks

5

Native scales kept separate

Artificial Analysis

Reference available

As of 2026-07-16

Benchmark-level evidence

Published Claude Opus 4.8 benchmark results

These are individual benchmarks, not collection rollups. Scores remain on their original scales, and agent or harness results stay labelled as system configurations.

How sources are reviewed →

Spring Prompt benchmark

BulletBench

304

60-second bullet Elo

Chess decision quality under a shared 60-second game clock and a 10+1 per-move clock.

Ladder Elo: anchored to a Stockfish skill ladder (random mover = 400). Internally consistent ordering, not FIDE-calibrated.

1 published configuration

Claude Opus 4.8 (low)

304

Bullet 95% interval
168–416
10+1 Elo
0
Bullet games
96
Bullet median move time
1.52s
Average bullet game cost
$0.0896
Source: Spring Prompt Recorded 3 Aug 2026 Ladder v1 · controls 10+1 and 60 Source record →

Spring Prompt benchmark

ROASBench

3 configurations

Average ROASBench score

A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.

3 published configurations

Anthropic: Claude Opus 4.8

9.54

Average contribution profit
$-312,597
Average ROAS
0.87×
Completed runs
3
Score variability
±0.96

Anthropic: Claude Opus 4.8 · High

10.41

Average contribution profit
$-352,542
Average ROAS
0.84×
Completed runs
3
Score variability
±2.28

Anthropic: Claude Opus 4.8 · Low

9.98

Average contribution profit
$-299,490
Average ROAS
0.89×
Completed runs
3
Score variability
±2.08
Source: Spring Prompt Recorded 25 Jul 2026 Published ROASBench cache · 12-month simulation Source record →

Spring Prompt benchmark

World Cup prediction benchmark

296

Tournament prediction points

Fixture-by-fixture football predictions frozen before the tournament and graded on outcomes, scorelines, goals, penalties, and cards.

32 fixtures graded from the frozen prediction artifact.

1 published configuration

Claude Opus 4.8

296

Graded matches
32
Exact-score points
15
Result points
36
Goal points
66
Source: Spring Prompt Recorded 20 Jul 2026 Predictions frozen 2026-06-11 Source record →

source-native benchmark

Coding Agent Configurations

2 configurations

Coding Agent Index v1.2 point score

Artificial Analysis Coding Agent Index v1.2 point scores and task-specific operational measurements for exact agent and model configurations.

These are exact agent-plus-model configurations, not bare-model scores.

2 published configurations

Claude Code - Opus 4.8 (max)

54.9

Mean cost per task
$7.70
Mean wall time
23.1 min

Claude Code - Opus 4.8 (medium)

48.9

Mean cost per task
$3.26
Mean wall time
12.4 min
Source: Artificial Analysis Coding Agent Index Recorded 20 Jul 2026 V1.2 Source record →

source-native benchmark

Repository Issue Resolution

83.8%

OpenHands SWE-Bench resolved

Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.

Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.

1 published configuration

OpenHands v1.24.0 + claude-opus-4-8

83.8%

95% task-bootstrap interval
80.6–87.0%
Task evidence
exact-headline-reconstruction
Source: OpenHands Index Recorded 30 Jun 2026 OpenHands · SWE-Bench 2026.06.30-3015ac6 Source record →

Independent operational reference

Artificial Analysis details

Representative configuration: Claude Opus 4.8 (Adaptive Reasoning, Max Effort). Lifecycle: active; released 2026-05-28.

Input price

$5.00 / 1M tokens

Output price

$25.00 / 1M tokens

Median output speed

54.8 tok/s

Operational reference only. Price and throughput are not model-quality scores and never affect Spring Prompt comparisons.

Early-user signal

Initial community opinions

Anecdotal · never scored

A paraphrased editorial synthesis of 4 Reddit discussions created in the 14 days after launch (28 May 2026 to 11 Jun 2026). Anecdotal context only — it never affects scores or rankings.

The search-indexed launch-window sample was sharply divided on Claude Opus 4.8. Some users reported more reliable tool use and longer stable engineering runs, while others encountered verbosity, slow inference, language errors, repeated actions, and rapid quota use. Experiences appeared to differ across languages, interfaces, effort settings, and task types.

Independent early-test signal

Early technical field tests

X · anecdotal · never scored

An editorial paraphrase of 2 launch-window field tests from 2 independent authors on X, including 2 reports with a described method or inspectable artifact. The window runs from 28 May 2026 up to 11 Jun 2026, and exact model identity was editor reviewed.

Early Opus 4.8 evidence combined an autonomous, long-horizon RPG build with two agent-benchmark results that moved in different directions. The artifact demonstrated broad planning, document creation, play-testing, and deployment, while the benchmark report found regressions on simulated business and planning tasks. The safest interpretation is high artifact-building ability with uneven agent performance by task and reasoning setting.

These reports are selectively surfaced and are not a representative sample. They never affect benchmark scores, rankings, winners, or comparison outcomes.

Business Skills V3 · proposed

How Spring Prompt plans to test Claude Opus 4.8

The setup below is proposed and may change until preflight and execution approval are complete. Existing benchmark evidence above does not authorize or stand in for a Business Skills V3 result.

Proposed reasoning
High reasoning
Provider revision
Exact provider revision will be resolved and frozen only after preflight and execution approval
Tools and service
No tools · Anthropic standard routing
Planning configuration
claude-opus-4.8-high
Generation controls
Provider-managed reasoning; no temperature override proposed

Future first-party coverage

15 proposed Business Skills V3 task areas

These links describe evaluation contracts, not published Claude Opus 4.8 results.

Model comparisons

Compare Claude Opus 4.8 side by side

Editorially reviewed comparisons appear first. Every page matches only benchmark records with the same reviewed protocol key and does not manufacture an overall winner.

Publication safeguards

What must pass before a V3 result appears

  1. 1First-party task-local comparisons and eligible external evidence must both be present.
  2. 2Model and provider configuration identity must match the reviewed release exactly.
  3. 3Coverage, reliability, judge diagnostics, and sealed stability checks must pass.
  4. 4Uncertainty and missing evidence remain visible when results are published.

Stable family URL

Evidence can grow without changing the page

New reviewed benchmark snapshots, operational facts, community themes, and early field-test syntheses can be added here while the canonical model-family identity remains fixed.

More from Anthropic

Other owned model profiles