What ARC-AGI-3 tests
ARC-AGI-3 evaluates an agent in unfamiliar, interactive environments. The system must explore, infer how an environment works, pursue a goal, and adapt its actions from feedback rather than answer a static question once.
That makes it useful evidence about interactive reasoning, but it also makes the evaluation setup part of the thing being measured.
How the score works
The primary method is Relative Human Action Efficiency (RHAE). Completion matters, and efficient completion matters: an evaluated system is compared with first-attempt human action baselines, then level and game results are aggregated onto a 0-100 scale.
A higher score is better, but a score is only interpretable alongside the exact environment set, scoring version, action policy, and budget used to produce it.
What a reported result actually measures
An ARC-AGI-3 result is a system result. It reflects the model plus the agent loop, prompts, state representation, tool access, action budget, retries, and scoring implementation. Calling it a bare model score hides variables that may materially change performance.
The useful comparison question is therefore: what did the evaluation hold constant, and what did it allow each system to vary?
Comparability policy
Why harness and state-management results stay separate
ARC describes its verified leaderboard as a no-domain-specific-harness comparison: evaluated systems receive the same ARC prompt and policy, while community harness research is reported separately. The technical report notes that useful harness improvements and progress on the official AGI-oriented leaderboard are evidence about different evaluation setups.
Springprompt therefore does not merge provider-managed state, custom memory, additional tools, or other agent-loop changes into the verified rows on this page. Those results can still be useful, but they should appear as a separate setup only after a canonical source and complete evaluation contract are available.
How Springprompt will compare results
Springprompt will display a result only with a complete evaluation contract and a reviewed source snapshot. Results produced under materially different harnesses, tools, budgets, sampling policies, or scoring protocols will be shown as distinct setups rather than merged into one model number.
The table contains all 26 scored entries on ARC's official verified ARC-AGI-3 track reviewed through 30 July 2026. They share ARC's official verified protocol, while each model family and reasoning variant remains a distinct configuration. Rows with no ARC-AGI-3 result and the separate community/custom-harness track are excluded rather than presented as comparable scores.