Benchmarks / Overall leaderboard

Combined by Spring Prompt

Overall leaderboard

Which current model does best across everything we publish? Each benchmark becomes a percentile among these models, then groups are weighted towards real business work. Hover a bar to see what it is made of. For launches and pace, see models over time.

Models ranked
18
Sources
6
Release
1 Oct 2026
Score
0–100, higher is better
Overall score from 0 to 100, higher is better, with each group's score
#ModelOverallSpring Prompt benchesWriting and preferenceFactualityGroups
1 Claude Opus 5.5Anthropic
78
7586–
2 Gemini 3.8 FlashGoogle
72
589188
3 GPT-6 AstraOpenAI
72
885550
4 GPT-6 SolOpenAI
67
863559
5 Muse Spark 1.3Meta
65
438894
6 Claude Fable 5.1Anthropic
62
4861100
7 Gemini 3.1 Pro PreviewGoogle
60
607838
8 DeepSeek V4.1 FlashDeepSeek
59
624769
9 Kimi K3Moonshot AI
57
387681
10 GPT-6 LunaOpenAI
48
732412
11 Grok 4.7xAI
43
70186
12 DeepSeek V4 Pro 0423DeepSeek
41
334259
13 Qwen3.8 Max (0902)Alibaba
40
225962
14 GLM 5.3Z.ai
35
126556
15 Gemini 3.5 Flash LiteGoogle
34
362938
16 Mistral Medium 3.5Mistral
23
33325
17 Claude Haiku 4.5Anthropic
13
16712
18 GLM 5V TurboZ.ai
12
51231

Not ranked yet, because they have results in only one group: Claude Sonnet 5.5, GPT-6.1 Sol.

How the score works

Spring Prompt benches 40%

Catalogue listings, ad budgets and fast chess: our own tests.

  • CatalogBench · 50% of the groupCatalogBench: reliably publish-ready; CatalogBench (sales brief): reliably publish-ready
  • ROASBench · 30% of the groupROASBench: overall score
  • BulletBench · 20% of the groupBulletBench Lightning 10+1: ladder Elo; BulletBench Bullet 60s: ladder Elo

Agents and tool use 25%

Multi-step tasks with tools: support desks, coding, function calls.

No source in this group has tested 5 or more of these models yet, so it does not count for now.

Writing and preference 20%

What people prefer in blind comparisons, and judged writing quality.

Factuality 15%

Sticking to the source when summarising, and factual answers.

  1. Same field for every number. Only the models in this table are compared, so a model is not lifted by a source that tested it against older ones. A benchmark counts once it has tested 5 of them.
  2. Percentiles, not raw scores. On each headline metric a model scores the share of the others it beats (ties count half), using its best configuration. Arena ratings, pass rates and Elo can then be averaged.
  3. Missing is not zero. A model needs CatalogBench or RoasBench and at least 2 groups. Groups it has no results in are left out and the others re-weighted; the dots show its coverage.
  4. What it does not show. Differences of a few points are within noise, and a percentile hides how far apart two models are. For a decision, read the benchmark pages; for your own product, test it on your data.

Derived from results reported by Arena (formerly LMArena) (Creative Commons Attribution 4.0 International), UGI Leaderboard (Apache License 2.0), Vectara Hallucination Leaderboard (Apache License 2.0), converted to percentiles by Spring Prompt.