The question we want to answer
Which models are best at tool calling and structured output?
A useful category rank must measure this outcome, not merely a broad intelligence score or a stylistically similar task.
Category ranking not yet published
Correct tool selection, arguments, schemas, values, multi-turn recovery, and explicit no-call behaviour under realistic constraints.
This page is not a model ranking.
We do not yet have enough defensible evidence to rank models for Tool Calling & Structured Output. No model order is shown here, and this category contributes nothing to the Spring Prompt overall leaderboard.
The question we want to answer
A useful category rank must measure this outcome, not merely a broad intelligence score or a stylistically similar task.
Why we are not ranking it yet
This V3 category is defined, but no model order is published until eligible external evidence and SpringPrompt pairwise evidence pass the coverage, independence, reliability, and stability gates.
Collection structure
Each skill has its own task plan and evidence status. A collection ranking cannot go live by silently substituting one facet for another.
Promotion gate
We do not fill these gaps with the retired internally judged scores. Saved outputs can still illustrate the task, but their old judge grades never determine a new rank.