The question we want to answer
Which models are best at software engineering and coding?
A useful category rank must measure this outcome, not merely a broad intelligence score or a stylistically similar task.
Category ranking not yet published
Repository-level implementation, debugging, testing, review, and maintenance with functional correctness and engineering quality kept distinct.
This page is not a model ranking.
We do not yet have enough defensible evidence to rank models for Software Engineering & Coding. No model order is shown here, and this category contributes nothing to the Spring Prompt overall leaderboard.
The question we want to answer
A useful category rank must measure this outcome, not merely a broad intelligence score or a stylistically similar task.
Why we are not ranking it yet
This V3 category is defined, but no model order is published until eligible external evidence and SpringPrompt pairwise evidence pass the coverage, independence, reliability, and stability gates.
Collection structure
Each skill has its own task plan and evidence status. A collection ranking cannot go live by silently substituting one facet for another.
Promotion gate
We do not fill these gaps with the retired internally judged scores. Saved outputs can still illustrate the task, but their old judge grades never determine a new rank.