Confirm Action

Are you sure you want to proceed?

Attributed source-native cohort explorer

STATE-Bench All workflows results

The publisher's aggregate across the evaluated workflow domains.

Exact configurations

9

Separate cohorts

3

Overall aggregate

Excluded

Each table is an independent comparison cohort. Row order is descending pass@1 within that cohort only; it is not a significance claim, confidence rank, cross-cohort rank, or cross-version leaderboard. The ± value is publisher-reported pass@1 standard deviation across runs. It is dispersion, not a standard error or confidence interval.

Main v0.8

2 configurations Harnessed models

Current v0.8 main-track submissions. Compare these two configurations only with one another.

Protocol 0.8.0 · source track main

Exact source configuration Pass@1 ± SD Pass5 Mean UX Benchmark cost/task Verification
GPT-5.5 OpenAI · harnessed model
Model
GPT-5.5
Reasoning
high
Agent
Not reported (source field blank)
Source entry
main-06
58.9% ± 0.6 SD run dispersion · not CI 44.2% 3.62source scale / 5 $0.1842publisher-reported Verified Submitted 2026-07-01
GPT-5.4 OpenAI · harnessed model
Model
GPT-5.4
Reasoning
high
Agent
Not reported (source field blank)
Source entry
main-07
58.6% ± 0.5 SD run dispersion · not CI 41.8% 3.67source scale / 5 $0.0874publisher-reported Verified Submitted 2026-07-01

Rows are shown from higher to lower All workflows pass@1 inside Main v0.8 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.

Main v0.7

5 configurations Harnessed models

Earlier v0.7 main-track submissions retained as their own historical comparison cohort.

Protocol 0.7.0 · source track main

Exact source configuration Pass@1 ± SD Pass5 Mean UX Benchmark cost/task Verification
GPT-5.4 OpenAI · harnessed model
Model
GPT-5.4
Reasoning
high
Agent
Not reported (source field blank)
Source entry
main-01
55.7% ± 1.9 SD run dispersion · not CI 38.0% 3.49source scale / 5 $0.0809publisher-reported Verified Submitted 2026-05-25
Claude Opus 4.7 Anthropic · harnessed model
Model
Claude Opus 4.7
Reasoning
high
Agent
Not reported (source field blank)
Source entry
main-02
53.4% ± 1.7 SD run dispersion · not CI 33.9% 3.46source scale / 5 $0.0773publisher-reported Verified Submitted 2026-05-29
Kimi-K2.6 Moonshot AI · harnessed model
Model
Kimi-K2.6
Reasoning
Not reported
Agent
Not reported (source field blank)
Source entry
main-03
48.3% ± 2.1 SD run dispersion · not CI 29.3% 3.37source scale / 5 $0.0496publisher-reported Verified Submitted 2026-05-25
DeepSeek-v4-Pro DeepSeek · harnessed model
Model
DeepSeek-v4-Pro
Reasoning
Not reported
Agent
Not reported (source field blank)
Source entry
main-04
47.2% ± 0.6 SD run dispersion · not CI 25.3% 3.36source scale / 5 Not reported Verified Submitted 2026-05-25
GPT-5.4 OpenAI · harnessed model
Model
GPT-5.4
Reasoning
default
Agent
Not reported (source field blank)
Source entry
main-05
46.9% ± 0.9 SD run dispersion · not CI 26.2% 3.41source scale / 5 $0.0351publisher-reported Verified Submitted 2026-05-25

Rows are shown from higher to lower All workflows pass@1 inside Main v0.7 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.

Agent Learning v0.4.4

2 configurations Complete agent systems

Memory/learning agent-system results. These rows are not bare foundation-model measurements.

Protocol 0.4.4 · source track memory

Exact source configuration Pass@1 ± SD Pass5 Mean UX Benchmark cost/task Verification
GPT 5.1 + Foundry Memory Microsoft Foundry · agent
Model
GPT 5.1 + Foundry Memory
Reasoning
Not reported
Agent
Not reported (source field blank)
Source entry
memory-01
58.3% ± 4.0 SD run dispersion · not CI 37.3% 3.93source scale / 5 $0.0360publisher-reported Verified Submitted 2026-06-08
GPT 5.1 OpenAI · agent
Model
GPT 5.1
Reasoning
no-memory
Agent
Not reported (source field blank)
Source entry
memory-02
53.3% ± 3.3 SD run dispersion · not CI 32.7% 3.87source scale / 5 $0.0310publisher-reported Verified Submitted 2026-06-08

Rows are shown from higher to lower All workflows pass@1 inside Agent Learning v0.4.4 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.

How to read these results

Pass@1 is the source-listed single-run completion percentage. Pass5, UX and benchmark cost per task are shown as separate source fields and do not modify pass@1.

The ± value is publisher-reported pass@1 standard deviation across runs. It is dispersion, not a standard error or confidence interval.

Not part of the overall leaderboard

STATE-Bench evaluates harnessed model or agent configurations on a source-native scale. Protocol versions and tracks are not mutually comparable, and the score has not been calibrated to SpringPrompt's cross-task predicted-fit scale.

No unified STATE-Bench rank · no Spring Prompt aggregate points

Frequently asked

Can I compare STATE-Bench v0.8 with v0.7?

No. Protocol versions are separate cohorts. This page never creates a combined order across v0.8, v0.7, or Agent Learning v0.4.4.

Does ± mean a confidence interval?

No. STATE-Bench reports pass@1 standard deviation across runs. It is run-to-run dispersion, not a standard error or confidence interval.

Are the memory-track rows bare-model results?

No. The v0.4.4 memory track evaluates complete agent-learning systems, so those rows remain labelled as agent configurations.

Does price or UX change the pass@1 order?

No. Rows are displayed by pass@1 within their exact cohort. Pass^5, UX and source benchmark cost are separate reported fields.