Confirm Action

Are you sure you want to proceed?

Attributed source-native cohort explorer

STATE-Bench Shopping assistant results

Shopping-assistant workflow tasks in STATE-Bench.

Exact configurations

9

Separate cohorts

3

Overall aggregate

Excluded

Each table is an independent comparison cohort. Row order is descending pass@1 within that cohort only; it is not a significance claim, confidence rank, cross-cohort rank, or cross-version leaderboard. The ± value is publisher-reported pass@1 standard deviation across runs. It is dispersion, not a standard error or confidence interval.

Main v0.8

2 configurations Harnessed models

Current v0.8 main-track submissions. Compare these two configurations only with one another.

Protocol 0.8.0 · source track main

Exact source configuration Pass@1 ± SD Pass5 Mean UX Benchmark cost/task Verification
GPT-5.4 OpenAI · harnessed model
Model
GPT-5.4
Reasoning
high
Agent
Not reported (source field blank)
Source entry
main-07
55.0% ± 1.0 SD run dispersion · not CI 45.0% 3.85source scale / 5 $0.0475publisher-reported Verified Submitted 2026-07-01
GPT-5.5 OpenAI · harnessed model
Model
GPT-5.5
Reasoning
high
Agent
Not reported (source field blank)
Source entry
main-06
55.0% ± 1.0 SD run dispersion · not CI 45.0% 3.87source scale / 5 $0.1170publisher-reported Verified Submitted 2026-07-01

Rows are shown from higher to lower Shopping assistant pass@1 inside Main v0.8 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.

Main v0.7

5 configurations Harnessed models

Earlier v0.7 main-track submissions retained as their own historical comparison cohort.

Protocol 0.7.0 · source track main

Exact source configuration Pass@1 ± SD Pass5 Mean UX Benchmark cost/task Verification
Claude Opus 4.7 Anthropic · harnessed model
Model
Claude Opus 4.7
Reasoning
high
Agent
Not reported (source field blank)
Source entry
main-02
54.0% ± 1.0 SD run dispersion · not CI 37.0% 3.72source scale / 5 $0.0496publisher-reported Verified Submitted 2026-05-29
GPT-5.4 OpenAI · harnessed model
Model
GPT-5.4
Reasoning
high
Agent
Not reported (source field blank)
Source entry
main-01
53.6% ± 2.1 SD run dispersion · not CI 40.0% 3.61source scale / 5 $0.0486publisher-reported Verified Submitted 2026-05-25
GPT-5.4 OpenAI · harnessed model
Model
GPT-5.4
Reasoning
default
Agent
Not reported (source field blank)
Source entry
main-05
50.3% ± 1.3 SD run dispersion · not CI 30.8% 3.55source scale / 5 $0.0216publisher-reported Verified Submitted 2026-05-25
DeepSeek-v4-Pro DeepSeek · harnessed model
Model
DeepSeek-v4-Pro
Reasoning
Not reported
Agent
Not reported (source field blank)
Source entry
main-04
48.4% ± 2.2 SD run dispersion · not CI 31.0% 3.54source scale / 5 Not reported Verified Submitted 2026-05-25
Kimi-K2.6 Moonshot AI · harnessed model
Model
Kimi-K2.6
Reasoning
Not reported
Agent
Not reported (source field blank)
Source entry
main-03
47.9% ± 2.2 SD run dispersion · not CI 36.0% 3.54source scale / 5 $0.0279publisher-reported Verified Submitted 2026-05-25

Rows are shown from higher to lower Shopping assistant pass@1 inside Main v0.7 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.

Agent Learning v0.4.4

2 configurations Complete agent systems

Memory/learning agent-system results. These rows are not bare foundation-model measurements.

Protocol 0.4.4 · source track memory

Exact source configuration Pass@1 ± SD Pass5 Mean UX Benchmark cost/task Verification
GPT 5.1 + Foundry Memory Microsoft Foundry · agent
Model
GPT 5.1 + Foundry Memory
Reasoning
Not reported
Agent
Not reported (source field blank)
Source entry
memory-01
57.0% ± 2.0 SD run dispersion · not CI 42.0% 3.74source scale / 5 $0.0160publisher-reported Verified Submitted 2026-06-08
GPT 5.1 OpenAI · agent
Model
GPT 5.1
Reasoning
no-memory
Agent
Not reported (source field blank)
Source entry
memory-02
53.0% ± 3.0 SD run dispersion · not CI 38.0% 3.81source scale / 5 $0.0140publisher-reported Verified Submitted 2026-06-08

Rows are shown from higher to lower Shopping assistant pass@1 inside Agent Learning v0.4.4 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.

How to read these results

Pass@1 is the source-listed single-run completion percentage. Pass5, UX and benchmark cost per task are shown as separate source fields and do not modify pass@1.

The ± value is publisher-reported pass@1 standard deviation across runs. It is dispersion, not a standard error or confidence interval.

Not part of the overall leaderboard

STATE-Bench evaluates harnessed model or agent configurations on a source-native scale. Protocol versions and tracks are not mutually comparable, and the score has not been calibrated to SpringPrompt's cross-task predicted-fit scale.

No unified STATE-Bench rank · no Spring Prompt aggregate points

Frequently asked

Can I compare STATE-Bench v0.8 with v0.7?

No. Protocol versions are separate cohorts. This page never creates a combined order across v0.8, v0.7, or Agent Learning v0.4.4.

Does ± mean a confidence interval?

No. STATE-Bench reports pass@1 standard deviation across runs. It is run-to-run dispersion, not a standard error or confidence interval.

Are the memory-track rows bare-model results?

No. The v0.4.4 memory track evaluates complete agent-learning systems, so those rows remain labelled as agent configurations.

Does price or UX change the pass@1 order?

No. Rows are displayed by pass@1 within their exact cohort. Pass^5, UX and source benchmark cost are separate reported fields.