Confirm Action

Are you sure you want to proceed?

Model evidence and rankings by task

We publish reviewed predicted-fit composites where several external signals must be combined, and attributed direct scores where a benchmark measures the target task itself.

The two score types stay separate. Direct source-native scores are excluded from the overall writing-fit standing unless scale compatibility is established. Uncertainty and operational coverage are disclosed per benchmark and are never assumed uniform.

The catalogue below covers all 15 core task categories. A status card is not a ranking: it shows why evidence is still insufficient or links to a narrower published signal, and it never enters the overall leaderboard.

Business & work

Customer Support Workflow Agents

Source-native STATE-Bench task-completion, user-experience and cost results for stateful enterprise workflows. Results are compared only inside the exact benchmark-version and track cohort.

9 exact configurations across 3 separate cohorts

Source-native values · no cross-cohort rank

Professional Knowledge Work & Strategy

Which models are best at professional knowledge work and strategy?

Category ranking not yet published

Not a ranking · excluded from overall

Business Workflow Automation

Which models are best at business workflow automation?

Category ranking not yet published

Not a ranking · excluded from overall

Customer Support & Service Resolution

Which models are best at customer support and service resolution?

Category ranking not yet published

Not a ranking · excluded from overall

Sales & Marketing

Which models are best at sales and marketing work?

Category ranking not yet published

Not a ranking · excluded from overall

Legal Document Work

Which models are best at legal document work?

Category ranking not yet published

Not a ranking · excluded from overall

Finance, Accounting & Spreadsheets

Which models are best at finance, accounting, and spreadsheet work?

Category ranking not yet published

Not a ranking · excluded from overall

Data Analysis & SQL

Which models are best at data analysis and SQL?

Category ranking not yet published

Not a ranking · excluded from overall

Research & Evidence Synthesis

Which models are best at research and evidence synthesis?

Category ranking not yet published

Not a ranking · excluded from overall

Business Writing & Email

Which models are best at business writing and email?

Category ranking not yet published

Not a ranking · excluded from overall

Document Comprehension & Extraction

Which models are best at document comprehension and extraction?

Category ranking not yet published

Not a ranking · excluded from overall

Summarization & Grounded Knowledge

Which models are best at summarization and grounded knowledge work?

Category ranking not yet published

Not a ranking · excluded from overall

Presentations, Charts & Visual Communication

Which models are best at presentations, charts, and visual communication?

Category ranking not yet published

Not a ranking · excluded from overall

Coding & agents

Spring Prompt planned model roster

Model profiles in our first-party collection

These 25 family-level profiles use stable URLs while the proposed Business Skills V3 evaluation moves through preflight and approval. Reasoning configurations stay inside the base model page rather than creating duplicate URLs.

This page is Spring Prompt, running in public

We just did this for every model. Do it for your prompt.

The rankings above come from running real tasks through real models and scoring every output. Spring Prompt is that same engine — pointed at your prompt, your test cases, and your definition of good.

  • Generate test cases from your prompt — no eval set required to start.
  • Compare models side by side with quality, cost and latency in one matrix.
  • Optimise the winner until the scores say it's ready to ship.
Experiment · Cold outreach email

Prompt × model results

12 test cases · 3 evals
Claude Opus
GPT-5
Gemini
v1
7.1
6.8
7.4
v2
8.3
7.9
8.0
v3
9.2
8.6
8.4
Best combo: v3 × Claude Opus
9.2 quality · $0.004/run · 1.8s