Confirm Action

Are you sure you want to proceed?

Methodology

Data sources and scoring

We use licensed and public benchmark evidence for two distinct publication types: task-specific predicted-fit composites and source-native direct benchmark outcomes. Their scales stay separate. Source-listed identities, known configurations, uncertainty treatment, missing data and attribution remain inspectable.

Active sources

EQ-Bench 3

Public benchmark signals used in the Spring Prompt predicted-fit beta.

Visit source ↗
Attribution
EQ-Bench 3 by Samuel Paech / EQ-bench
Source snapshot fetched
2026-07-16
Configurations
33

Source methodology ↗ Source terms ↗

The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.

EQ-Bench Creative Writing v3

Public benchmark signals used in the Spring Prompt predicted-fit beta.

Visit source ↗
Attribution
Creative Writing v3 by Samuel Paech / EQ-bench
Source snapshot fetched
2026-07-16
Configurations
33

Source methodology ↗ Source terms ↗

The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.

LMArena Leaderboard Dataset

Public benchmark signals used in the Spring Prompt predicted-fit beta.

Visit source ↗
Attribution
Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0
Source snapshot fetched
2026-07-16
Configurations
33

Source methodology ↗ Source terms ↗

The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.

UGI Leaderboard

Public benchmark signals used in the Spring Prompt predicted-fit beta.

Visit source ↗
Attribution
UGI Leaderboard by DontPlanToEnd, via Hugging Face Spaces
Source snapshot fetched
2026-07-16
Configurations
33

Source methodology ↗ Source terms ↗

The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.

Structured Output Benchmark (SOB)

Direct structured-output outcome score and uncertainty rank range.

Visit source ↗
Attribution
Structured Output Benchmark (SOB) by Interfaze / JigsawStack, Inc. (MIT License)
Source release date
2026-07-02
Configurations
35

Source methodology ↗ Source terms ↗

The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.

OpenHands Index

Official OpenHands configuration scores and task-sidecar uncertainty availability.

Visit source ↗
Attribution
OpenHands Index by the OpenHands contributors (Apache-2.0)
Source snapshot fetched
2026-07-17
Configurations
34

The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.

Microsoft STATE-Bench

Source-native workflow completion, run-to-run dispersion, UX and benchmark cost within exact STATE-Bench cohorts.

Visit source ↗
Attribution
STATE-Bench by Microsoft and STATE-Bench contributors (MIT License)
Source snapshot fetched
2026-07-20
Configurations
9

Source methodology ↗ Source terms ↗

The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.

Artificial Analysis Coding Agent Index

Coding Agent Index v1.2 component outcomes, point score, mean benchmark cost per task and mean wall time per task.

Visit source ↗
Attribution
Coding Agent Index data sourced from Artificial Analysis
Source snapshot fetched
2026-07-20
Configurations
44

Source methodology ↗ Source terms ↗

The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.

Artificial Analysis

Operational price and output-throughput metadata only; not a predicted-fit input.

Visit source ↗
Attribution
Operational price and performance data sourced from Artificial Analysis
Source snapshot fetched
2026-07-16
Configurations
22

Source methodology ↗

Verification and current model availability are not exposed by this API.

The benchmark numbers are source-listed or deterministically derived from the attributed release. Spring Prompt attribution and interpretation do not imply that a source endorses our task mappings, rankings or operational overlays.

Model identity: conservative by design

  1. Every source-listed model variant is retained, even when Spring Prompt cannot run it.
  2. A source-listed variant may link to our runnable catalog only through an authoritative identifier that is unique on both sides or a human-reviewed alias decision. Those are different evidence paths and are labelled separately.
  3. Provider, revision, reasoning effort, and service tier are score-affecting identity fields when the source supplies them; published values cannot be discarded.
  4. Name similarity creates a review suggestion only. It never makes a result rankable and never selects the highest-scoring variant.

Published task rankings

We publish only reviewed task mappings or direct benchmarks with a defensible target match. Unsupported categories receive no score.

Predictive public beta: Email Writing and Business Writing only. These two tasks are non-strict predictions assembled from several attributed benchmark sources, and they are the only task scores entering the overall Writing-fit standing. They are not task-success percentages or fresh direct model runs.

Business Writing

2026-07-16-business-candidate-v6

Predicted ability to produce clear, coherent and audience-appropriate business prose such as briefs, updates, proposals and memos.

Coverage floor
55%
Signal groups
3+
Lineages
3+

Email Writing

2026-07-16-business-candidate-v6

Predicted ability to write accurate, audience-aware, appropriately toned and action-oriented business, sales and support emails.

Coverage floor
55%
Signal groups
3+
Lineages
3+
Direct benchmark publications retain their native source contracts. Structured Output publishes model-family outcomes with its reviewed marginal rank-range method. Repository Issue Resolution publishes exact OpenHands configurations and mixed task-bootstrap score-interval availability. STATE-Bench publishes exact configurations in hard-partitioned protocol cohorts. None is blended with the predictive Writing tasks or enters the overall Writing-fit standing.

Structured Output Reliability

Structured Output Benchmark Overall · 35 model families

Directly measured ability to return accurate values in the requested structured schema across the benchmark's evaluated text, image and audio modalities.

100 × evaluated-modality coverage × equal-weight mean of the seven publisher-defined unified component outcomes.

Overall exclusion: The source-native direct score is not on the same calibrated scale as SpringPrompt predicted-fit task scores; cross-task scale compatibility has not been established.

Rank ranges use overlap of marginal record-cluster bootstrap intervals for Overall; simultaneous rank confidence is not claimed.

Repository Issue Resolution

OpenHands SWE-Bench resolved · 34 exact agent configurations

Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.

Official OpenHands SWE-Bench resolved percentage for the exact OpenHands and language-model configuration.

Overall exclusion: The publisher-native resolved percentage measures an OpenHands configuration on SWE-Bench. It is neither a calibrated SpringPrompt task-fit score nor a bare-model score, so cross-task scale compatibility has not been established.

Task-bootstrap score intervals are available for 29 configurations and explicitly unavailable for 5. The official source point order covers all rows; no full-cohort rank confidence is claimed.

Coding Agent Configurations

Coding Agent Index v1.2 point score · 44 exact agent configurations

Artificial Analysis Coding Agent Index v1.2 point scores and task-specific operational measurements for exact agent and model configurations.

Simple average of pass@1 across DeepSWE, Terminal-Bench v2 and SWE-Atlas-QnA for rows with all three components.

Overall exclusion: The source score evaluates complete coding-agent and model configurations on the Artificial Analysis v1.2 protocol. It is neither a bare-model score nor calibrated to SpringPrompt's cross-task predicted-fit scale.

Only 40 rows with all three v1.2 components enter the point, cost or wall-time orders. 4 retained two-component legacy rows are shown explicitly unranked. Values are source point estimates; no interval, significance or rank confidence is inferred.

Customer Support Workflow Agents

STATE-Bench source-native fields · 9 exact configurations · 3 separate cohorts

Source-native STATE-Bench completion, UX and cost results for stateful enterprise workflows.

Overall exclusion: Protocol versions and the agent-learning track are not mutually comparable, and their source-native pass rates are not calibrated to Spring Prompt predicted fit.

Pass@1 standard deviation is publisher-reported run-to-run dispersion, not a standard error or confidence interval. No cross-cohort rank or significance claim is published.

Published ranking safeguards