Methodology
Data sources and scoring
We use licensed and public benchmark evidence for two distinct publication types: task-specific predicted-fit composites and source-native direct benchmark outcomes. Their scales stay separate. Source-listed identities, known configurations, uncertainty treatment, missing data and attribution remain inspectable.
Active sources
EQ-Bench 3
Public benchmark signals used in the Spring Prompt predicted-fit beta.
- Attribution
- EQ-Bench 3 by Samuel Paech / EQ-bench
- Source snapshot fetched
- 2026-07-16
- Configurations
- 33
Source methodology ↗ Source terms ↗
The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.
EQ-Bench Creative Writing v3
Public benchmark signals used in the Spring Prompt predicted-fit beta.
- Attribution
- Creative Writing v3 by Samuel Paech / EQ-bench
- Source snapshot fetched
- 2026-07-16
- Configurations
- 33
Source methodology ↗ Source terms ↗
The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.
LMArena Leaderboard Dataset
Public benchmark signals used in the Spring Prompt predicted-fit beta.
- Attribution
- Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0
- Source snapshot fetched
- 2026-07-16
- Configurations
- 33
Source methodology ↗ Source terms ↗
The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.
UGI Leaderboard
Public benchmark signals used in the Spring Prompt predicted-fit beta.
- Attribution
- UGI Leaderboard by DontPlanToEnd, via Hugging Face Spaces
- Source snapshot fetched
- 2026-07-16
- Configurations
- 33
Source methodology ↗ Source terms ↗
The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.
Structured Output Benchmark (SOB)
Direct structured-output outcome score and uncertainty rank range.
- Attribution
- Structured Output Benchmark (SOB) by Interfaze / JigsawStack, Inc. (MIT License)
- Source release date
- 2026-07-02
- Configurations
- 35
Source methodology ↗ Source terms ↗
The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.
OpenHands Index
Official OpenHands configuration scores and task-sidecar uncertainty availability.
- Attribution
- OpenHands Index by the OpenHands contributors (Apache-2.0)
- Source snapshot fetched
- 2026-07-17
- Configurations
- 34
The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.
Microsoft STATE-Bench
Source-native workflow completion, run-to-run dispersion, UX and benchmark cost within exact STATE-Bench cohorts.
- Attribution
- STATE-Bench by Microsoft and STATE-Bench contributors (MIT License)
- Source snapshot fetched
- 2026-07-20
- Configurations
- 9
Source methodology ↗ Source terms ↗
The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.
Artificial Analysis Coding Agent Index
Coding Agent Index v1.2 component outcomes, point score, mean benchmark cost per task and mean wall time per task.
- Attribution
- Coding Agent Index data sourced from Artificial Analysis
- Source snapshot fetched
- 2026-07-20
- Configurations
- 44
Source methodology ↗ Source terms ↗
The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.
Artificial Analysis
Operational price and output-throughput metadata only; not a predicted-fit input.
- Attribution
- Operational price and performance data sourced from Artificial Analysis
- Source snapshot fetched
- 2026-07-16
- Configurations
- 22
Verification and current model availability are not exposed by this API.
The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.
Model identity: conservative by design
- Every source-listed model variant is retained, even when Spring Prompt cannot run it.
- A source-listed variant may link to our runnable catalog only through an authoritative identifier that is unique on both sides or a human-reviewed alias decision. Those are different evidence paths and are labelled separately.
- Provider, revision, reasoning effort, and service tier are score-affecting identity fields when the source supplies them; published values cannot be discarded.
- Name similarity creates a review suggestion only. It never makes a result rankable and never selects the highest-scoring variant.
Published task rankings
We publish only reviewed task mappings or direct benchmarks with a defensible target match. Unsupported categories receive no score.
Business Writing
2026-07-16-business-candidate-v6
Predicted ability to produce clear, coherent and audience-appropriate business prose such as briefs, updates, proposals and memos.
- Coverage floor
- 55%
- Signal groups
- 3+
- Lineages
- 3+
Email Writing
2026-07-16-business-candidate-v6
Predicted ability to write accurate, audience-aware, appropriately toned and action-oriented business, sales and support emails.
- Coverage floor
- 55%
- Signal groups
- 3+
- Lineages
- 3+
Structured Output Reliability
Structured Output Benchmark Overall · 35 model families
Directly measured ability to return accurate values in the requested structured schema across the benchmark's evaluated text, image and audio modalities.
100 × evaluated-modality coverage × equal-weight mean of the seven publisher-defined unified component outcomes.
Overall exclusion: The source-native direct score is not on the same calibrated scale as SpringPrompt predicted-fit task scores; cross-task scale compatibility has not been established.
Rank ranges use overlap of marginal record-cluster bootstrap intervals for Overall; simultaneous rank confidence is not claimed.
Repository Issue Resolution
OpenHands SWE-Bench resolved · 34 exact agent configurations
Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.
Official OpenHands SWE-Bench resolved percentage for the exact OpenHands and language-model configuration.
Overall exclusion: The publisher-native resolved percentage measures an OpenHands configuration on SWE-Bench. It is neither a calibrated SpringPrompt task-fit score nor a bare-model score, so cross-task scale compatibility has not been established.
Task-bootstrap score intervals are available for 29 configurations and explicitly unavailable for 5. The official source point order covers all rows; no full-cohort rank confidence is claimed.
Coding Agent Configurations
Coding Agent Index v1.2 point score · 44 exact agent configurations
Artificial Analysis Coding Agent Index v1.2 point scores and task-specific operational measurements for exact agent and model configurations.
Simple average of pass@1 across DeepSWE, Terminal-Bench v2 and SWE-Atlas-QnA for rows with all three components.
Overall exclusion: The source score evaluates complete coding-agent and model configurations on the Artificial Analysis v1.2 protocol. It is neither a bare-model score nor calibrated to SpringPrompt's cross-task predicted-fit scale.
Only 40 rows with all three v1.2 components enter the point, cost or wall-time orders. 4 retained two-component legacy rows are shown explicitly unranked. Values are source point estimates; no interval, significance or rank confidence is inferred.
Customer Support Workflow Agents
STATE-Bench source-native fields · 9 exact configurations · 3 separate cohorts
Source-native STATE-Bench completion, UX and cost results for stateful enterprise workflows.
Overall exclusion: Protocol versions and the agent-learning track are not mutually comparable, and their source-native pass rates are not calibrated to Spring Prompt predicted fit.
Pass@1 standard deviation is publisher-reported run-to-run dispersion, not a standard error or confidence interval. No cross-cohort rank or significance claim is published.
Published ranking safeguards
- Predicted fit is a relative within-task estimate, not a task-success percentage, probability, guarantee, or cross-category score.
- Evidence coverage is displayed separately from performance; missing signals are not treated as model failures.
- Only model families meeting the published coverage, independent-signal-group, and evidence-lineage floors enter the provisional order.
- The order is explicitly non-strict. Small score differences should be treated as the same performance band.
- Product-family identity matching does not prove that a model is runnable through Spring Prompt, so no runnable model link is invented.
- Artificial Analysis price and output-throughput metadata is joined only by an exact reviewed family key, remains visibly tied to its representative configuration, and never changes predicted fit.
- Structured Output retains the publisher-defined Overall formula and source-native 0–100 scale; it is never relabelled as predicted fit or a task-success guarantee.
- Its 95% Overall intervals use record-cluster bootstrap resampling with component covariance preserved. Overlap produces a marginal rank range, not simultaneous confidence for the full rank vector.
- Structured Output is explicitly excluded from the overall Writing-fit standing because scale compatibility with predicted-fit scores has not been established.
- Repository Issue Resolution preserves the exact OpenHands agent version and source model identity. It never projects those outcomes onto a bare language model.
- Its 29 exact task sidecars receive marginal score intervals; the five affected sidecars receive no inferred interval, and no uniform full-cohort rank confidence is claimed.
- Artificial Analysis Coding Agent Index v1.2 positions only complete three-component agent configurations. Its four two-component legacy rows stay visible but receive no quality, cheapest or fastest position.
- Coding Agent cost per task and wall time per task are source-native operational measurements for the same exact configuration; they contribute no quality points.
- STATE-Bench Main v0.8, Main v0.7 and Agent Learning v0.4.4 are hard partitions. No combined order is computed across versions or tracks.
- STATE-Bench displays exact model, reasoning, agent, verification and submission identity. Its reported pass@1 SD is run-to-run dispersion—not a standard error, confidence interval or rank range.
- ROASBench and BulletBench remain Spring Prompt-run focused benchmarks and are labelled separately from external-source composites.