Model evidence and rankings by task
We publish reviewed predicted-fit composites where several external signals must be combined, and attributed direct scores where a benchmark measures the target task itself.
The two score types stay separate. Direct source-native scores are excluded from the overall writing-fit standing unless scale compatibility is established. Uncertainty and operational coverage are disclosed per benchmark and are never assumed uniform.
The catalogue below covers all 15 core task categories. A status card is not a ranking: it shows why evidence is still insufficient or links to a narrower published signal, and it never enters the overall leaderboard.
Business & work
Customer Support Workflow Agents
Source-native STATE-Bench task-completion, user-experience and cost results for stateful enterprise workflows. Results are compared only inside the exact benchmark-version and track cohort.
Source-native values · no cross-cohort rank
Professional Knowledge Work & Strategy
Which models are best at professional knowledge work and strategy?
Category ranking not yet published
Not a ranking · excluded from overall
Business Workflow Automation
Which models are best at business workflow automation?
Category ranking not yet published
Not a ranking · excluded from overall
Customer Support & Service Resolution
Which models are best at customer support and service resolution?
Category ranking not yet published
Not a ranking · excluded from overall
Sales & Marketing
Which models are best at sales and marketing work?
Category ranking not yet published
Not a ranking · excluded from overall
Legal Document Work
Which models are best at legal document work?
Category ranking not yet published
Not a ranking · excluded from overall
Finance, Accounting & Spreadsheets
Which models are best at finance, accounting, and spreadsheet work?
Category ranking not yet published
Not a ranking · excluded from overall
Data Analysis & SQL
Which models are best at data analysis and SQL?
Category ranking not yet published
Not a ranking · excluded from overall
Research & Evidence Synthesis
Which models are best at research and evidence synthesis?
Category ranking not yet published
Not a ranking · excluded from overall
Business Writing & Email
Which models are best at business writing and email?
Category ranking not yet published
Not a ranking · excluded from overall
Document Comprehension & Extraction
Which models are best at document comprehension and extraction?
Category ranking not yet published
Not a ranking · excluded from overall
Summarization & Grounded Knowledge
Which models are best at summarization and grounded knowledge work?
Category ranking not yet published
Not a ranking · excluded from overall
Presentations, Charts & Visual Communication
Which models are best at presentations, charts, and visual communication?
Category ranking not yet published
Not a ranking · excluded from overall
Coding & agents
Repository Issue Resolution
Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.
Official source score · exact agent identity
Software Engineering & Coding
Which models are best at software engineering and coding?
Category ranking not yet published
Not a ranking · excluded from overall
Frontend, UI & Web Apps
Which models are best at frontend, UI, and web-app creation?
Category ranking not yet published
Not a ranking · excluded from overall
Tool Calling & Structured Output
Which models are best at tool calling and structured output?
Category ranking not yet published
Not a ranking · excluded from overall
Spring Prompt planned model roster
Model profiles in our first-party collection
These 25 family-level profiles use stable URLs while the proposed Business Skills V3 evaluation moves through preflight and approval. Reasoning configurations stay inside the base model page rather than creating duplicate URLs.
Anthropic
DeepSeek
Mistral
Moonshot AI
OpenAI
xAI
Z.ai
This page is Spring Prompt, running in public
We just did this for every model. Do it for your prompt.
The rankings above come from running real tasks through real models and scoring every output. Spring Prompt is that same engine — pointed at your prompt, your test cases, and your definition of good.
- Generate test cases from your prompt — no eval set required to start.
- Compare models side by side with quality, cost and latency in one matrix.
- Optimise the winner until the scores say it's ready to ship.
Prompt × model results
12 test cases · 3 evals