Benchmarks / Overall leaderboard
Overall leaderboard
Which current model does best across everything we publish? Each benchmark becomes a percentile among these models, then groups are weighted towards real business work. Hover a bar to see what it is made of. For launches and pace, see models over time.
- Models ranked
- 18
- Sources
- 6
- Release
- 1 Oct 2026
- Score
- 0–100, higher is better
- Spring Prompt benches 40%
- Agents and tool use not counted yet
- Writing and preference 20%
- Factuality 15%
| # | Model | Overall | Spring Prompt benches | Writing and preference | Factuality | Groups |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5Anthropic | 78 | 75 | 86 | – | |
| 2 | Gemini 3.8 FlashGoogle | 72 | 58 | 91 | 88 | |
| 3 | GPT-6 AstraOpenAI | 72 | 88 | 55 | 50 | |
| 4 | GPT-6 SolOpenAI | 67 | 86 | 35 | 59 | |
| 5 | Muse Spark 1.3Meta | 65 | 43 | 88 | 94 | |
| 6 | Claude Fable 5.1Anthropic | 62 | 48 | 61 | 100 | |
| 7 | Gemini 3.1 Pro PreviewGoogle | 60 | 60 | 78 | 38 | |
| 8 | DeepSeek V4.1 FlashDeepSeek | 59 | 62 | 47 | 69 | |
| 9 | Kimi K3Moonshot AI | 57 | 38 | 76 | 81 | |
| 10 | GPT-6 LunaOpenAI | 48 | 73 | 24 | 12 | |
| 11 | Grok 4.7xAI | 43 | 70 | 18 | 6 | |
| 12 | DeepSeek V4 Pro 0423DeepSeek | 41 | 33 | 42 | 59 | |
| 13 | Qwen3.8 Max (0902)Alibaba | 40 | 22 | 59 | 62 | |
| 14 | GLM 5.3Z.ai | 35 | 12 | 65 | 56 | |
| 15 | Gemini 3.5 Flash LiteGoogle | 34 | 36 | 29 | 38 | |
| 16 | Mistral Medium 3.5Mistral | 23 | 33 | 3 | 25 | |
| 17 | Claude Haiku 4.5Anthropic | 13 | 16 | 7 | 12 | |
| 18 | GLM 5V TurboZ.ai | 12 | 5 | 12 | 31 |
Not ranked yet, because they have results in only one group: Claude Sonnet 5.5, GPT-6.1 Sol.
How the score works
Spring Prompt benches 40%
Catalogue listings, ad budgets and fast chess: our own tests.
- CatalogBench · 50% of the groupCatalogBench: reliably publish-ready; CatalogBench (sales brief): reliably publish-ready
- ROASBench · 30% of the groupROASBench: overall score
- BulletBench · 20% of the groupBulletBench Lightning 10+1: ladder Elo; BulletBench Bullet 60s: ladder Elo
Agents and tool use 25%
Multi-step tasks with tools: support desks, coding, function calls.
- tau2-benchtau2 v1.0.1, retail: task success (pass^1); tau2 v1.0.1, airline: task success (pass^1); tau2 v1.0.1, telecom: task success (pass^1)
- Berkeley Function Calling Leaderboard (BFCL) V4BFCL: overall accuracy
- OpenHands IndexOpenHands Index: average
- Microsoft STATE-BenchSTATE-Bench v0.7: task success (pass@1)
No source in this group has tested 5 or more of these models yet, so it does not count for now.
Writing and preference 20%
What people prefer in blind comparisons, and judged writing quality.
- Arena textArena Text: overall
- UGI LeaderboardUGI: writing score
Factuality 15%
Sticking to the source when summarising, and factual answers.
- Vectara Hallucination LeaderboardVectara: hallucination rate
- Arena factualityArena Text factuality: overall
- Same field for every number. Only the models in this table are compared, so a model is not lifted by a source that tested it against older ones. A benchmark counts once it has tested 5 of them.
- Percentiles, not raw scores. On each headline metric a model scores the share of the others it beats (ties count half), using its best configuration. Arena ratings, pass rates and Elo can then be averaged.
- Missing is not zero. A model needs CatalogBench or RoasBench and at least 2 groups. Groups it has no results in are left out and the others re-weighted; the dots show its coverage.
- What it does not show. Differences of a few points are within noise, and a percentile hides how far apart two models are. For a decision, read the benchmark pages; for your own product, test it on your data.
Derived from results reported by Arena (formerly LMArena) (Creative Commons Attribution 4.0 International), UGI Leaderboard (Apache License 2.0), Vectara Hallucination Leaderboard (Apache License 2.0), converted to percentiles by Spring Prompt.