Attributed source-native cohort explorer
STATE-Bench Shopping assistant results
Shopping-assistant workflow tasks in STATE-Bench.
Exact configurations
9
Separate cohorts
3
Overall aggregate
Excluded
Each table is an independent comparison cohort. Row order is descending pass@1 within that cohort only; it is not a significance claim, confidence rank, cross-cohort rank, or cross-version leaderboard. The ± value is publisher-reported pass@1 standard deviation across runs. It is dispersion, not a standard error or confidence interval.
Main v0.8
2 configurations Harnessed modelsCurrent v0.8 main-track submissions. Compare these two configurations only with one another.
Protocol 0.8.0 · source track main
| Exact source configuration | Pass@1 ± SD | Pass5 | Mean UX | Benchmark cost/task | Verification |
|---|---|---|---|---|---|
GPT-5.4
OpenAI · harnessed model
|
55.0% ± 1.0 SD run dispersion · not CI | 45.0% | 3.85source scale / 5 | $0.0475publisher-reported | Verified Submitted 2026-07-01 |
GPT-5.5
OpenAI · harnessed model
|
55.0% ± 1.0 SD run dispersion · not CI | 45.0% | 3.87source scale / 5 | $0.1170publisher-reported | Verified Submitted 2026-07-01 |
Rows are shown from higher to lower Shopping assistant pass@1 inside Main v0.8 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.
Main v0.7
5 configurations Harnessed modelsEarlier v0.7 main-track submissions retained as their own historical comparison cohort.
Protocol 0.7.0 · source track main
| Exact source configuration | Pass@1 ± SD | Pass5 | Mean UX | Benchmark cost/task | Verification |
|---|---|---|---|---|---|
Claude Opus 4.7
Anthropic · harnessed model
|
54.0% ± 1.0 SD run dispersion · not CI | 37.0% | 3.72source scale / 5 | $0.0496publisher-reported | Verified Submitted 2026-05-29 |
GPT-5.4
OpenAI · harnessed model
|
53.6% ± 2.1 SD run dispersion · not CI | 40.0% | 3.61source scale / 5 | $0.0486publisher-reported | Verified Submitted 2026-05-25 |
GPT-5.4
OpenAI · harnessed model
|
50.3% ± 1.3 SD run dispersion · not CI | 30.8% | 3.55source scale / 5 | $0.0216publisher-reported | Verified Submitted 2026-05-25 |
DeepSeek-v4-Pro
DeepSeek · harnessed model
|
48.4% ± 2.2 SD run dispersion · not CI | 31.0% | 3.54source scale / 5 | Not reported | Verified Submitted 2026-05-25 |
Kimi-K2.6
Moonshot AI · harnessed model
|
47.9% ± 2.2 SD run dispersion · not CI | 36.0% | 3.54source scale / 5 | $0.0279publisher-reported | Verified Submitted 2026-05-25 |
Rows are shown from higher to lower Shopping assistant pass@1 inside Main v0.7 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.
Agent Learning v0.4.4
2 configurations Complete agent systemsMemory/learning agent-system results. These rows are not bare foundation-model measurements.
Protocol 0.4.4 · source track memory
| Exact source configuration | Pass@1 ± SD | Pass5 | Mean UX | Benchmark cost/task | Verification |
|---|---|---|---|---|---|
GPT 5.1 + Foundry Memory
Microsoft Foundry · agent
|
57.0% ± 2.0 SD run dispersion · not CI | 42.0% | 3.74source scale / 5 | $0.0160publisher-reported | Verified Submitted 2026-06-08 |
GPT 5.1
OpenAI · agent
|
53.0% ± 3.0 SD run dispersion · not CI | 38.0% | 3.81source scale / 5 | $0.0140publisher-reported | Verified Submitted 2026-06-08 |
Rows are shown from higher to lower Shopping assistant pass@1 inside Agent Learning v0.4.4 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.
How to read these results
Pass@1 is the source-listed single-run completion percentage. Pass5, UX and benchmark cost per task are shown as separate source fields and do not modify pass@1.
The ± value is publisher-reported pass@1 standard deviation across runs. It is dispersion, not a standard error or confidence interval.
Not part of the overall leaderboard
STATE-Bench evaluates harnessed model or agent configurations on a source-native scale. Protocol versions and tracks are not mutually comparable, and the score has not been calibrated to SpringPrompt's cross-task predicted-fit scale.
No unified STATE-Bench rank · no Spring Prompt aggregate points
Frequently asked
Can I compare STATE-Bench v0.8 with v0.7?
No. Protocol versions are separate cohorts. This page never creates a combined order across v0.8, v0.7, or Agent Learning v0.4.4.
Does ± mean a confidence interval?
No. STATE-Bench reports pass@1 standard deviation across runs. It is run-to-run dispersion, not a standard error or confidence interval.
Are the memory-track rows bare-model results?
No. The v0.4.4 memory track evaluates complete agent-learning systems, so those rows remain labelled as agent configurations.
Does price or UX change the pass@1 order?
No. Rows are displayed by pass@1 within their exact cohort. Pass^5, UX and source benchmark cost are separate reported fields.