Confirm Action

Are you sure you want to proceed?

ARC Prize Foundation · ARC-AGI-3 (2026) · Opus 4.7

Verified results through 2026-04-16

Opus 4.7 on ARC-AGI-3: scores and setup

ARC-AGI-3 is an interactive reasoning benchmark. A system must explore unfamiliar game-like environments, infer their rules, pursue a goal, and act efficiently without being given task instructions.

Best published configuration
Opus 4.7 · High
Best score
0.18%
Primary metric
RHAE · 0–100
Evaluation subject
model configuration in ARC's evaluation loop

ARC Prize verified · semi-private

Opus 4.7 scores on ARC-AGI-3

1 configuration · 1 model family

This model view includes every scored Opus 4.7 configuration in RHAE · semi-private. Return to the benchmark page to compare all published rows. Scores are RHAE percentages, not question-answer accuracy.

Opus 4.7 published configurations

Each mark is one scored reasoning configuration published for this model on the same benchmark source and evaluation contract. Labels contain the exact RHAE score.

Compare all models →
  1. High 0.18%
0.0 0.1 0.2

RHAE · 0–100

Official verified semi-private track only; entries with no ARC-AGI-3 score are omitted. Community/custom-harness track ↗

How to interpret the result

What do the Opus 4.7 results mean?

This view isolates one model family without changing the evaluation track or pretending its configurations are equal-compute runs.

1. Read the best configuration

The strongest published Opus 4.7 setup scored 0.18% on ARC-AGI-3 score at High reasoning. That is a system result under the cited evaluation protocol.

2. Reasoning settings matter

Changing a stated reasoning variant or other configuration detail changes the evaluated system. The spread between rows is evidence about the whole setup, not noise to average into one bare score.

3. Compare on the same track

Use the main benchmark page for comparison with other rows from the same source field and evaluation contract. Results from a different harness remain a separate methodology.

Best takeaway: Opus 4.7's published ARC-AGI-3 score depends on its stated configuration and the evaluation contract.

Treat the Opus 4.7 rows as configuration results under the cited protocol, not as a context-free model rating.

Every row comes from RHAE · semi-private, recorded 16 Apr 2026.

#ModelConfigurationVerified score
1 Opus 4.7Best High 0.18%
How these setups compareShared protocol, reported differences, and explicit unknowns Open contract
Comparison track
ARC Prize official verified leaderboard, using its semi-private environment set and RHAE scoring.
Harness and prompts
No domain-specific external harness; ARC reports the same benchmark prompt and evaluation policy across providers.
What varies
Model family and stated reasoning effort. Provider internals are not controlled or fully disclosed by ARC.
Tools
No model-facing external tools. Provider-side implementation details remain a black box.
Budget
ARC action policy is shared; exact per-result token, runtime, and compute budgets are not published on the cited result pages.
Comparison limit
Comparable as official leaderboard entries, but not as equal-compute experiments.

What ARC-AGI-3 tests

ARC-AGI-3 evaluates an agent in unfamiliar, interactive environments. The system must explore, infer how an environment works, pursue a goal, and adapt its actions from feedback rather than answer a static question once.

That makes it useful evidence about interactive reasoning, but it also makes the evaluation setup part of the thing being measured.

ARC Prize replay viewer showing GPT-5.6 Sol playing the ft09 ARC-AGI-3 environment
A frame from ARC Prize's public ft09 replay for GPT-5.6 Sol at Max reasoning. The replay makes the benchmark's multi-turn environment, actions, and carried reasoning visible. Open the interactive ARC Prize replay ↗

How the score works

The primary method is Relative Human Action Efficiency (RHAE). Completion matters, and efficient completion matters: an evaluated system is compared with first-attempt human action baselines, then level and game results are aggregated onto a 0-100 scale.

A higher score is better, but a score is only interpretable alongside the exact environment set, scoring version, action policy, and budget used to produce it.

What a reported result actually measures

An ARC-AGI-3 result is a system result. It reflects the model plus the agent loop, prompts, state representation, tool access, action budget, retries, and scoring implementation. Calling it a bare model score hides variables that may materially change performance.

The useful comparison question is therefore: what did the evaluation hold constant, and what did it allow each system to vary?

Comparability policy

Why harness and state-management results stay separate

ARC describes its verified leaderboard as a no-domain-specific-harness comparison: evaluated systems receive the same ARC prompt and policy, while community harness research is reported separately. The technical report notes that useful harness improvements and progress on the official AGI-oriented leaderboard are evidence about different evaluation setups.

Springprompt therefore does not merge provider-managed state, custom memory, additional tools, or other agent-loop changes into the verified rows on this page. Those results can still be useful, but they should appear as a separate setup only after a canonical source and complete evaluation contract are available.

How Springprompt will compare results

Springprompt will display a result only with a complete evaluation contract and a reviewed source snapshot. Results produced under materially different harnesses, tools, budgets, sampling policies, or scoring protocols will be shown as distinct setups rather than merged into one model number.

This filtered table contains only the published Opus 4.7 configurations from the complete official comparison. It does not hide other model families from the benchmark page or merge results from a different evaluation track.

Official ARC resources 10 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.ARC Prize 2026 competition overview ↗ARC Prize Foundation · what-it-tests · retrieved 2026-07-30 · evidence 5054a50b2715
  2. 2.ARC-AGI-3 scoring methodology ↗ARC Prize Foundation · scoring · retrieved 2026-07-30 · evidence 9d9c0e1f82f1
  3. 3.ARC-AGI-3 technical report ↗ARC Prize Foundation · what-it-tests · retrieved 2026-07-30 · evidence 418e98f08f36
  4. 4.ARC-AGI-3 human dataset ↗ARC Prize Foundation · scoring · retrieved 2026-07-30 · evidence 9f203a410d37
  5. 5.ARC-AGI-3 agent benchmarking guide ↗ARC Prize Foundation · system-not-model · retrieved 2026-07-30 · evidence b9f4f3971da2
  6. 6.ARC Prize verified GPT-5.6 series results ↗ARC Prize Foundation · comparison-policy · retrieved 2026-07-30 · evidence b989d728be8c
  7. 7.ARC Prize verified Claude Opus 5 results ↗ARC Prize Foundation · comparison-policy · retrieved 2026-07-30 · evidence 6b7f36186144
  8. 8.ARC Prize verified leaderboard ↗ARC Prize Foundation · comparison-policy · retrieved 2026-07-30 · evidence 7d60ca68fda1
  9. 9.ARC Prize community leaderboard ↗ARC Prize Foundation · current-harness-point · retrieved 2026-07-30 · evidence 4378996ca7dd
  10. 10.ARC-AGI open-source toolkit ↗ARC Prize Foundation · official-links · retrieved 2026-07-30 · evidence 8d2d9e023b7b
  11. 11.ARC-AGI-3 official benchmarking agent repository ↗ARC Prize Foundation · system-not-model · retrieved 2026-07-30 · evidence 0b68f14bad72
  12. 12.ARC-AGI-3 agents and harness examples ↗ARC Prize Foundation · current-harness-point · retrieved 2026-07-30 · evidence 7a61f3095e57
  13. 13.Official ft09 GPT-5.6 Sol replay ↗ARC Prize Foundation · what-it-tests · retrieved 2026-07-30 · evidence d064b591ef43

Read the official scoring methodology ↗