Attributed source-native cohort explorer
STATE-Bench Travel results
Travel-planning and booking workflow tasks in STATE-Bench.
Exact configurations
9
Separate cohorts
3
Overall aggregate
Excluded
Each table is an independent comparison cohort. Row order is descending pass@1 within that cohort only; it is not a significance claim, confidence rank, cross-cohort rank, or cross-version leaderboard. The ± value is publisher-reported pass@1 standard deviation across runs. It is dispersion, not a standard error or confidence interval.
Main v0.8
2 configurations Harnessed modelsCurrent v0.8 main-track submissions. Compare these two configurations only with one another.
Protocol 0.8.0 · source track main
| Exact source configuration | Pass@1 ± SD | Pass5 | Mean UX | Benchmark cost/task | Verification |
|---|---|---|---|---|---|
GPT-5.4
OpenAI · harnessed model
|
62.0% ± 1.0 SD run dispersion · not CI | 39.0% | 3.58source scale / 5 | $0.1339publisher-reported | Verified Submitted 2026-07-01 |
GPT-5.5
OpenAI · harnessed model
|
62.0% ± 2.0 SD run dispersion · not CI | 45.0% | 3.61source scale / 5 | $0.2802publisher-reported | Verified Submitted 2026-07-01 |
Rows are shown from higher to lower Travel pass@1 inside Main v0.8 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.
Main v0.7
5 configurations Harnessed modelsEarlier v0.7 main-track submissions retained as their own historical comparison cohort.
Protocol 0.7.0 · source track main
| Exact source configuration | Pass@1 ± SD | Pass5 | Mean UX | Benchmark cost/task | Verification |
|---|---|---|---|---|---|
GPT-5.4
OpenAI · harnessed model
|
55.9% ± 2.8 SD run dispersion · not CI | 36.0% | 3.37source scale / 5 | $0.1224publisher-reported | Verified Submitted 2026-05-25 |
Claude Opus 4.7
Anthropic · harnessed model
|
55.0% ± 1.0 SD run dispersion · not CI | 36.0% | 3.47source scale / 5 | $0.1068publisher-reported | Verified Submitted 2026-05-29 |
Kimi-K2.6
Moonshot AI · harnessed model
|
51.9% ± 4.0 SD run dispersion · not CI | 26.0% | 3.22source scale / 5 | $0.0871publisher-reported | Verified Submitted 2026-05-25 |
DeepSeek-v4-Pro
DeepSeek · harnessed model
|
47.6% ± 2.7 SD run dispersion · not CI | 22.0% | 3.04source scale / 5 | Not reported | Verified Submitted 2026-05-25 |
GPT-5.4
OpenAI · harnessed model
|
43.2% ± 1.5 SD run dispersion · not CI | 22.9% | 3.19source scale / 5 | $0.0565publisher-reported | Verified Submitted 2026-05-25 |
Rows are shown from higher to lower Travel pass@1 inside Main v0.7 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.
Agent Learning v0.4.4
2 configurations Complete agent systemsMemory/learning agent-system results. These rows are not bare foundation-model measurements.
Protocol 0.4.4 · source track memory
| Exact source configuration | Pass@1 ± SD | Pass5 | Mean UX | Benchmark cost/task | Verification |
|---|---|---|---|---|---|
GPT 5.1 + Foundry Memory
Microsoft Foundry · agent
|
60.0% ± 6.0 SD run dispersion · not CI | 34.0% | 4.03source scale / 5 | $0.0650publisher-reported | Verified Submitted 2026-06-08 |
GPT 5.1
OpenAI · agent
|
56.0% ± 5.0 SD run dispersion · not CI | 30.0% | 3.96source scale / 5 | $0.0610publisher-reported | Verified Submitted 2026-06-08 |
Rows are shown from higher to lower Travel pass@1 inside Agent Learning v0.4.4 only. The order uses no cross-cohort normalization and makes no statistical-significance or confidence-rank claim.
How to read these results
Pass@1 is the source-listed single-run completion percentage. Pass5, UX and benchmark cost per task are shown as separate source fields and do not modify pass@1.
The ± value is publisher-reported pass@1 standard deviation across runs. It is dispersion, not a standard error or confidence interval.
Not part of the overall leaderboard
STATE-Bench evaluates harnessed model or agent configurations on a source-native scale. Protocol versions and tracks are not mutually comparable, and the score has not been calibrated to SpringPrompt's cross-task predicted-fit scale.
No unified STATE-Bench rank · no Spring Prompt aggregate points
Frequently asked
Can I compare STATE-Bench v0.8 with v0.7?
No. Protocol versions are separate cohorts. This page never creates a combined order across v0.8, v0.7, or Agent Learning v0.4.4.
Does ± mean a confidence interval?
No. STATE-Bench reports pass@1 standard deviation across runs. It is run-to-run dispersion, not a standard error or confidence interval.
Are the memory-track rows bare-model results?
No. The v0.4.4 memory track evaluates complete agent-learning systems, so those rows remain labelled as agent configurations.
Does price or UX change the pass@1 order?
No. Rows are displayed by pass@1 within their exact cohort. Pass^5, UX and source benchmark cost are separate reported fields.