Attributed source-native cohort explorer
STATE-Bench Customer support results
Customer-support workflow tasks in STATE-Bench.
Exact configurations
9
Separate cohorts
3
Overall aggregate
Excluded
Each table is an independent comparison cohort. Row order is descending pass@1 within that cohort only; it is not a significance claim, confidence rank, cross-cohort rank, or cross-version leaderboard. The ± value is publisher-reported pass@1 standard deviation across runs. It is dispersion, not a standard error or confidence interval.
Main v0.8
2 configurations Harnessed modelsCurrent v0.8 main-track submissions. Compare these two configurations only with one another.
Protocol 0.8.0 · source track main
| Exact source configuration | Pass@1 ± SD | Pass5 | Mean UX | Benchmark cost/task | Verification |
|---|---|---|---|---|---|
GPT-5.4
OpenAI · harnessed model
|
59.0% ± 2.0 SD run dispersion · not CI | 42.0% | 3.59source scale / 5 | $0.0808publisher-reported | Verified Submitted 2026-07-01 |
GPT-5.5
OpenAI · harnessed model
|
59.0% ± 1.0 SD run dispersion · not CI | 43.0% | 3.38source scale / 5 | $0.1553publisher-reported | Verified Submitted 2026-07-01 |
Rows are shown from higher to lower Customer support pass@1 inside Main v0.8 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.
Main v0.7
5 configurations Harnessed modelsEarlier v0.7 main-track submissions retained as their own historical comparison cohort.
Protocol 0.7.0 · source track main
| Exact source configuration | Pass@1 ± SD | Pass5 | Mean UX | Benchmark cost/task | Verification |
|---|---|---|---|---|---|
GPT-5.4
OpenAI · harnessed model
|
57.6% ± 2.5 SD run dispersion · not CI | 38.0% | 3.49source scale / 5 | $0.0716publisher-reported | Verified Submitted 2026-05-25 |
Claude Opus 4.7
Anthropic · harnessed model
|
51.0% ± 3.0 SD run dispersion · not CI | 28.0% | 3.19source scale / 5 | $0.0755publisher-reported | Verified Submitted 2026-05-29 |
GPT-5.4
OpenAI · harnessed model
|
47.2% ± 2.1 SD run dispersion · not CI | 28.1% | 3.49source scale / 5 | $0.0271publisher-reported | Verified Submitted 2026-05-25 |
DeepSeek-v4-Pro
DeepSeek · harnessed model
|
45.6% ± 1.4 SD run dispersion · not CI | 23.0% | 3.50source scale / 5 | Not reported | Verified Submitted 2026-05-25 |
Kimi-K2.6
Moonshot AI · harnessed model
|
45.1% ± 1.8 SD run dispersion · not CI | 26.0% | 3.35source scale / 5 | $0.0339publisher-reported | Verified Submitted 2026-05-25 |
Rows are shown from higher to lower Customer support pass@1 inside Main v0.7 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.
Agent Learning v0.4.4
2 configurations Complete agent systemsMemory/learning agent-system results. These rows are not bare foundation-model measurements.
Protocol 0.4.4 · source track memory
| Exact source configuration | Pass@1 ± SD | Pass5 | Mean UX | Benchmark cost/task | Verification |
|---|---|---|---|---|---|
GPT 5.1 + Foundry Memory
Microsoft Foundry · agent
|
58.0% ± 4.0 SD run dispersion · not CI | 36.0% | 4.03source scale / 5 | $0.0280publisher-reported | Verified Submitted 2026-06-08 |
GPT 5.1
OpenAI · agent
|
51.0% ± 2.0 SD run dispersion · not CI | 30.0% | 3.84source scale / 5 | $0.0170publisher-reported | Verified Submitted 2026-06-08 |
Rows are shown from higher to lower Customer support pass@1 inside Agent Learning v0.4.4 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.
How to read these results
Pass@1 is the source-listed single-run completion percentage. Pass5, UX and benchmark cost per task are shown as separate source fields and do not modify pass@1.
The ± value is publisher-reported pass@1 standard deviation across runs. It is dispersion, not a standard error or confidence interval.
Not part of the overall leaderboard
STATE-Bench evaluates harnessed model or agent configurations on a source-native scale. Protocol versions and tracks are not mutually comparable, and the score has not been calibrated to SpringPrompt's cross-task predicted-fit scale.
No unified STATE-Bench rank · no Spring Prompt aggregate points
Frequently asked
Can I compare STATE-Bench v0.8 with v0.7?
No. Protocol versions are separate cohorts. This page never creates a combined order across v0.8, v0.7, or Agent Learning v0.4.4.
Does ± mean a confidence interval?
No. STATE-Bench reports pass@1 standard deviation across runs. It is run-to-run dispersion, not a standard error or confidence interval.
Are the memory-track rows bare-model results?
No. The v0.4.4 memory track evaluates complete agent-learning systems, so those rows remain labelled as agent configurations.
Does price or UX change the pass@1 order?
No. Rows are displayed by pass@1 within their exact cohort. Pass^5, UX and source benchmark cost are separate reported fields.