Benchmarks / UGI Leaderboard

Reported by UGI Leaderboard

UGI Leaderboard

UGI's blend of intelligence, style, repetition and length adherence in writing, tuned to average human preference.

Last updated 2 Oct 2026

Results dated
6 Sep 2025 to 2 Oct 2026
Results
379 configurations of 230 models
Unit
score out of 100
Licence
Apache License 2.0

Writing score: GPT-5.6 Sol

Top 15 of 230 results · score out of 100, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 Gemini 3.8 Flash (medium reasoning)Google 78.6
  2. 2 Gemini 3.7 Flash (medium reasoning)Google 77.5
  3. 3 Claude Fable 5 (high reasoning)Anthropic 74.2
  4. 4 Gemini 3.5 Flash (medium reasoning)Google 72.5
  5. 5 Gemini 3.1 Pro Preview (low reasoning)Google 72.2
  6. 6 Claude 3.7 Sonnet (no reasoning)Anthropic 71.2
  7. 7 Claude Opus 4.6 (high reasoning)Anthropic 70.9
  8. 8 Claude Opus 4.5 (no reasoning)Anthropic 70.4
  9. 9 Claude Opus 5.5 (low reasoning)Anthropic 70.0
  10. 10 Claude Opus 4.1 (no reasoning)Anthropic 69.9
  11. 11 Gemini 3.6 Flash (medium reasoning)Google 69.8
  12. 12 GPT-5.5 (extra-high reasoning)OpenAI 69.7
  13. 13 GPT-6 Astra (high reasoning)OpenAI 68.9
  14. 14 GPT-5.4 (high reasoning)OpenAI 68.8
  15. 15 DeepSeek-V4-Pro (0423, reasoning on)DeepSeek 68.4
  16. 27 GPT-5.6 Sol (medium reasoning)OpenAI 65.4

Full results

UGI: writing score, score out of 100, higher is better
#ModelWriting score
score out of 100, higher is better
Price
$ per million tokens, in / out
1 Gemini 3.8 Flash (medium reasoning)Google · best of 3 settings
78.6
$1.50 / $7.50
2 Gemini 3.7 Flash (medium reasoning)Google · best of 3 settings
77.5
$1.50 / $7.50
3 Claude Fable 5 (high reasoning)Anthropic · best of 5 settings
74.2
$10 / $50
4 Gemini 3.5 Flash (medium reasoning)Google · best of 4 settings
72.5
$1.50 / $9
5 Gemini 3.1 Pro Preview (low reasoning)Google · best of 3 settings
72.2
$2 / $12
6 Claude 3.7 Sonnet (no reasoning)Anthropic · best of 2 settings
71.2
–
7 Claude Opus 4.6 (high reasoning)Anthropic · best of 4 settings
70.9
$5 / $25
8 Claude Opus 4.5 (no reasoning)Anthropic · best of 2 settings
70.4
$5 / $25
9 Claude Opus 5.5 (low reasoning)Anthropic · best of 5 settings
70.0
$4 / $20
10 Claude Opus 4.1 (no reasoning)Anthropic · best of 2 settings
69.9
$15 / $75
11 Gemini 3.6 Flash (medium reasoning)Google · best of 3 settings
69.8
$1.50 / $7.50
12 GPT-5.5 (extra-high reasoning)OpenAI · best of 5 settings
69.7
$5 / $30
13 GPT-6 Astra (high reasoning)OpenAI · best of 4 settings
68.9
$10 / $50
14 GPT-5.4 (high reasoning)OpenAI · best of 5 settings
68.8
$2.50 / $15
15 DeepSeek-V4-Pro (0423, reasoning on)DeepSeek · best of 2 settings
68.4
$1.42 / $2.83
16 Claude Opus 4 (no reasoning)Anthropic · best of 2 settings
68.2
–
17 GPT-6.1 Sol (high reasoning)OpenAI · best of 4 settings
67.7
$2 / $10
17 Gemini 3 Flash Preview (medium reasoning)Google · best of 3 settings
67.7
$0.50 / $3
19 GPT-4.1OpenAI
67.5
$2 / $8
20 ChatGPT-4o (2025-03-26)OpenAI
67.4
–
21 GLM-5.2 (reasoning on)Z.ai · best of 2 settings
67.0
$1.40 / $4.40
21 Grok 4 (0709)xAI
67.0
–
23 Claude Sonnet 4.5 (no reasoning)Anthropic · best of 2 settings
66.9
$3 / $15
24 o3 (medium reasoning)OpenAI · best of 3 settings
66.8
$2 / $8
25 Gemini 3 Pro Preview (high reasoning)Google · best of 2 settings
66.7
–
26 Claude Opus 4.8 (max reasoning)Anthropic · best of 5 settings
65.9
$5 / $25
27 GPT-5.6 Sol (medium reasoning)OpenAI · best of 5 settings
65.4
$4 / $20
28 Gemini 2.5 ProGoogle
65.2
$1.25 / $10
29 Claude Sonnet 4.6 (high reasoning)Anthropic · best of 4 settings
64.6
$3 / $15
30 Kimi K2.6 (reasoning on)Moonshot AI · best of 2 settings
64.4
$0.95 / $4
31 Claude Fable 5.1 (medium reasoning)Anthropic · best of 3 settings
63.9
$10 / $50
32 Claude Sonnet 4 (no reasoning)Anthropic · best of 2 settings
63.7
$3 / $15
33 GPT-4o (2024-05-13)OpenAI
63.4
$5 / $15
34 Grok 4.6xAI
63.3
$2 / $6
35 Grok 4.20 Multi-Agent Beta (0309, 4 agents)xAI
63.1
–
36 Claude Sonnet 5 (max reasoning)Anthropic · best of 5 settings
62.7
$2 / $10
36 GPT-5 Chat (latest)OpenAI
62.7
–
38 Kimi K2.5 (reasoning on)Moonshot AI · best of 2 settings
62.4
$0.57 / $2.85
39 GLM-5.1 (reasoning on)Z.ai · best of 2 settings
61.7
$1.38 / $4.40
40 Claude Opus 4.7 (max reasoning)Anthropic · best of 4 settings
61.4
$5 / $25
41 o1 (low reasoning)OpenAI · best of 2 settings
61.0
$15 / $60
42 GPT-5.2 ChatOpenAI
60.8
–
43 GPT-5.3 ChatOpenAI
60.7
–
44 GPT-5.6 Terra (high reasoning)OpenAI · best of 5 settings
60.4
$2 / $12
45 GPT-5.1 ChatOpenAI
60.3
–
46 Grok 4.5xAI
60.0
$2 / $6
47 Kimi K2 ThinkingMoonshot AI
59.3
$0.60 / $2.50
48 Grok 3xAI
58.7
–
49 Qwen3.6-Plus (reasoning on)Alibaba · best of 2 settings
58.2
$0.33 / $1.95
50 Grok 4.3xAI
57.7
$1.25 / $2.50
51 MiMo-V2.5-Pro (reasoning on)Xiaomi · best of 2 settings
57.4
$0.43 / $0.87
51 Grok 4.20 Beta (0309, reasoning)xAI
57.4
–
53 GLM-4.5 (reasoning on)Z.ai · best of 2 settings
57.3
$0.60 / $2.20
53 DeepSeek-V3.2 (reasoning on)DeepSeek · best of 2 settings
57.3
$0.30 / $0.96
55 MiMo-V2.5 (reasoning on)Xiaomi · best of 2 settings
56.4
$0.17 / $0.34
56 GPT-5.1 (high reasoning)OpenAI · best of 3 settings
55.8
$1.25 / $10
57 GPT-6 Sol (high reasoning)OpenAI · best of 4 settings
55.6
$2 / $10
58 Grok 4.20 (0309, reasoning)xAI
55.3
–
59 GLM-5 (reasoning on)Z.ai · best of 2 settings
55.0
$0.95 / $2.55
60 DeepSeek-V4-Flash (0423, no reasoning)DeepSeek · best of 2 settings
54.6
$0.14 / $0.28
61 Kimi K2 (0905)Moonshot AI
54.4
$0.60 / $2.50
62 Claude 3 OpusAnthropic
54.3
–
63 Grok 4.7xAI
54.1
$2 / $6
64 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek · best of 2 settings
54.0
$0.27 / $1
64 DeepSeek-V3.2-Exp (reasoning on)DeepSeek · best of 2 settings
54.0
$0.27 / $0.41
64 DeepSeek-V3.2-SpecialeDeepSeek
54.0
–
67 GPT-5.6 Luna (high reasoning)OpenAI · best of 5 settings
53.7
$0.20 / $1.20
68 GPT-5 (high reasoning)OpenAI · best of 2 settings
53.5
$1.25 / $10
69 Grok 4 Fast (reasoning)xAI
53.4
–
70 GLM-4.6 (reasoning on)Z.ai · best of 2 settings
53.3
$0.50 / $2
71 Kimi K2 (0711)Moonshot AI
52.5
$0.57 / $2.30
72 Claude Haiku 4.5 (no reasoning)Anthropic · best of 2 settings
52.2
$1 / $5
73 DeepSeek-V3.1 (reasoning on)DeepSeek · best of 2 settings
51.9
$0.55 / $1.65
74 DeepSeek-R1DeepSeek
51.8
$0.70 / $2.50
75 Grok 4.1 Fast (reasoning)xAI
51.3
–
76 Hy3 Preview (reasoning on)Tencent · best of 2 settings
51.1
$0.18 / $0.60
77 Qwen3.5-397B-A17B (no reasoning)Alibaba
50.1
$0.55 / $3.50
78 GPT-5.2 (high reasoning)OpenAI · best of 4 settings
49.4
$1.75 / $14
79 DeepSeek-V3-0324DeepSeek
49.1
$0.25 / $1
80 Qwen3-235B-A22B-Instruct-2507Alibaba
49.0
$0.15 / $0.75
81 GLM-4.7 (reasoning on)Z.ai · best of 2 settings
47.4
$0.54 / $1.98
82 Grok 4.20 (0309, non-reasoning)xAI
47.2
–
83 GLM-4.5-Air (no reasoning)Z.ai · best of 2 settings
47.0
$0.14 / $0.86
84 o4-mini (low reasoning)OpenAI · best of 3 settings
46.9
$1.10 / $4.40
85 Step 3.5 FlashStepFun
45.9
$0.10 / $0.30
86 Mistral Medium 3.5 (no reasoning)Mistral AI · best of 2 settings
45.5
$1.50 / $7.50
87 Qwen3.6-35B-A3B (reasoning by prefill)Alibaba · best of 2 settings
45.4
$0.10 / $1
88 Gemma 3 27BGoogle
45.0
$0.12 / $0.20
89 Gemini 2.5 Flash Preview (09-2025, no reasoning)Google · best of 2 settings
44.9
–
89 MiMo-V2-Flash (reasoning on)Xiaomi · best of 2 settings
44.9
–
91 Grok 4.20 Beta (0309, non-reasoning)xAI
43.9
–
92 Gemma 4 26B A4B (reasoning by prefill)Google · best of 2 settings
43.8
$0.10 / $0.30
93 DeepSeek-V3DeepSeek
43.2
$0.26 / $1.03
94 Gemma 4 31B (reasoning by prefill)Google · best of 2 settings
43.1
$0.14 / $0.40
95 Grok 4 Fast (non-reasoning)xAI
43.0
–
95 Qwen3-VL-235B-A22B-InstructAlibaba
43.0
$0.30 / $1.50
97 Qwen3.5-35B-A3B (reasoning by prefill)Alibaba · best of 2 settings
42.6
$0.16 / $1.30
98 Qwen3.6-27B (reasoning by prefill)Alibaba · best of 2 settings
42.5
$0.30 / $3.20
99 Qwen3.5-27B (reasoning by prefill)Alibaba · best of 2 settings
42.4
$0.27 / $2.16
100 DeepSeek-R1-0528DeepSeek
42.3
$0.50 / $2.18
101 Grok 4.1 Fast (non-reasoning)xAI
41.9
–
102 Qwen3-Next-80B-A3B-InstructAlibaba
41.8
$0.10 / $1.10
103 MiniMax-M2.5MiniMax
41.7
$0.30 / $1.20
104 MiniMax-M2.1MiniMax
41.6
$0.30 / $1.20
104 Mistral Large 3 (no reasoning)Mistral AI · best of 2 settings
41.6
$0.50 / $1.50
106 Muse Glimmer 30B (extra-high reasoning)Meta · best of 4 settings
41.0
$0.30 / $1.20
107 Mistral Large 2 (2407)Mistral AI
40.9
$2 / $6
108 Trinity Large PreviewArcee AI
40.7
–
109 Command ACohere
40.6
$2.50 / $10
110 Mistral Small 4 (no reasoning)Mistral AI · best of 2 settings
40.3
$0.15 / $0.60
111 GLM-4.6V (reasoning on)Z.ai · best of 2 settings
39.7
$0.30 / $0.90
112 Qwen3.5-122B-A10B (no reasoning)Alibaba · best of 2 settings
39.5
$0.26 / $2.08
112 Qwen3.5-9B (reasoning by prefill)Alibaba · best of 2 settings
39.5
$0.10 / $0.15
112 Mistral Medium 3.1Mistral AI
39.5
$0.40 / $2
115 Qwen3-MaxAlibaba
38.6
$0.78 / $3.90
116 gpt-oss-120b (medium reasoning)OpenAI
38.5
$0.15 / $0.60
117 Jamba Large 1.7AI21 Labs
38.1
–
118 Mistral Large 2.1 (2411)Mistral AI
37.8
–
119 Kimi Linear 48B A3B InstructMoonshot AI
37.7
–
120 MiniMax-M2.7MiniMax
37.6
$0.30 / $1.20
121 Gemma 4 12B (reasoning by prefill)Google · best of 2 settings
36.3
–
121 Mistral Small 3.2 24BMistral AI
36.3
$0.094 / $0.25
123 MedGemma 27B TextGoogle
36.1
–
124 Qwen2.5-VL-72B-InstructAlibaba
35.6
$0.80 / $1
125 Gemma 2 27BGoogle
35.0
$0.65 / $0.65
125 Qwen3-VL-32B-InstructAlibaba
35.0
$0.10 / $0.42
127 QwQ-32BAlibaba
34.9
–
128 Qwen3-14B (no reasoning)Alibaba · best of 2 settings
34.8
$0.12 / $0.24
129 Qwen2.5-32B-InstructAlibaba
34.2
–
130 Magistral Small 1.2Mistral AI · best of 2 settings
34.0
–
131 Mistral Small 3Mistral AI
33.9
$0.05 / $0.08
132 Nemotron Nano 12B v2 VL (BF16)NVIDIA · best of 2 settings
33.3
–
133 Mistral NemoMistral AI
33.1
$0.023 / $0.03
134 Qwen3-32B (no reasoning)Alibaba · best of 2 settings
33.0
$0.14 / $0.40
135 Llama 3.1 70B InstructMeta
32.8
$0.40 / $0.40
136 Mistral Medium 3Mistral AI
32.6
$0.40 / $2
136 Mistral Small (2409)Mistral AI
32.6
–
138 Ring-1TAnt Group
32.4
–
138 Qwen3-235B-A22B-Thinking-2507Alibaba
32.4
$0.30 / $3
140 Qwen3-VL-235B-A22B-ThinkingAlibaba
32.3
$0.40 / $4
140 Command R+ (08-2024)Cohere
32.3
$2.50 / $10
142 Mistral Small 3.1 24BMistral AI
32.1
$0.35 / $0.56
143 Seed-OSS-36B-Instruct (512-token reasoning budget)ByteDance · best of 3 settings
32.0
–
144 Claude 3 HaikuAnthropic
31.9
–
145 Qwen3-Coder-30B-A3B-InstructAlibaba
31.8
$0.07 / $0.28
146 Devstral Small 2Mistral AI
31.6
–
147 Qwen2.5-72B-InstructAlibaba
31.5
$0.36 / $0.40
148 Ling-1TAnt Group
31.4
–
148 Ministral 3 14BMistral AI · best of 2 settings
31.4
$0.20 / $0.20
150 Qwen3.5-4B (reasoning by prefill)Alibaba · best of 2 settings
30.8
–
151 Qwen3-VL-8B-InstructAlibaba
30.3
$0.12 / $0.46
152 Qwen3-30B-A3B (no reasoning)Alibaba · best of 2 settings
30.2
$0.12 / $0.50
153 Apriel Nemotron 15B ThinkerServiceNow
30.0
–
154 Qwen3-4B-Instruct-2507Alibaba
29.9
–
154 Llama 3.1 405B InstructMeta
29.9
–
154 Gemma 3 12BGoogle
29.9
$0.05 / $0.15
157 Qwen2.5-14B-InstructAlibaba
29.8
–
158 Qwen2.5-7B-InstructAlibaba
29.7
$0.10 / $0.20
159 Qwen3.5-2B (no reasoning)Alibaba · best of 2 settings
29.0
–
160 Ministral 3 8BMistral AI · best of 2 settings
28.8
$0.15 / $0.15
161 Olmo 3 32B ThinkAi2
28.5
–
162 Qwen3-8B (no reasoning)Alibaba · best of 2 settings
28.0
$0.12 / $0.46
163 Qwen3-30B-A3B-Instruct-2507Alibaba
27.9
$0.09 / $0.30
164 LFM2-24B-A2BLiquid AI
27.6
–
165 Mixtral 8x22B InstructMistral AI
27.4
$2 / $6
166 Ministral 3 8B ReasoningMistral AI · best of 3 settings
27.2
–
166 Qwen3-Next-80B-A3B-ThinkingAlibaba
27.2
$0.15 / $1.20
168 InternLM3 8B InstructShanghai AI Lab
27.1
–
169 Llama 4 ScoutMeta
26.7
$0.18 / $0.59
170 Command R (08-2024)Cohere
26.5
$0.15 / $0.60
171 Llama 3.3 70B InstructMeta
26.2
$0.59 / $0.79
172 Olmo 3 7B ThinkAi2
26.0
–
173 Nemotron 3 Nano 30B A3BNVIDIA · best of 2 settings
25.9
$0.05 / $0.20
174 Falcon-H1 7B InstructTII
25.7
–
174 Phi-4Microsoft
25.7
$0.07 / $0.14
174 Mixtral 8x7B Instruct v0.1Mistral AI
25.7
–
177 Ministral 3 14B ReasoningMistral AI · best of 2 settings
25.6
–
177 EXAONE 4.0 32BLG AI Research
25.6
–
179 GLM-4.7-Flash (no reasoning)Z.ai · best of 2 settings
25.0
$0.06 / $0.40
179 Gemma 2 9BGoogle
25.0
–
181 Llama 3.1 8B InstructMeta
24.8
$0.05 / $0.08
182 gpt-oss-20b (low reasoning)OpenAI · best of 3 settings
24.6
$0.03 / $0.15
183 Kimi-VL-A3B-InstructMoonshot AI
24.4
–
183 Olmo 3 7B InstructAi2
24.4
–
185 Gemma 3 4BGoogle
24.0
$0.05 / $0.10
185 Llama 4 MaverickMeta
24.0
$0.27 / $0.85
187 Jamba Mini 1.7AI21 Labs
23.8
–
187 Reka Flash 3Reka AI
23.8
$0.10 / $0.20
189 Solar 10.7B Instruct v1.0Upstage
23.6
–
190 Nova 2 Lite (no reasoning)Amazon · best of 2 settings
23.5
$0.30 / $2.50
191 Qwen3.5-0.8B (no reasoning)Alibaba
23.0
–
192 Qwen3-VL-32B-ThinkingAlibaba
22.9
–
193 Falcon-H1 3B InstructTII
22.8
–
194 Qwen3-4B (no reasoning)Alibaba · best of 2 settings
22.4
–
195 Qwen2.5-1.5B-InstructAlibaba
22.2
–
196 Qwen2.5-VL-32B-InstructAlibaba
21.7
–
197 Gemma 4 E4B (reasoning by prefill)Google · best of 2 settings
21.6
–
198 Falcon-H1 1.5B InstructTII
21.0
–
199 Qwen3-30B-A3B-Thinking-2507Alibaba
20.6
$0.20 / $2.40
199 Kimi-VL-A3B-Thinking (2506)Moonshot AI
20.6
–
201 Qwen2.5-VL-3B-InstructAlibaba
20.4
–
202 Gemma 2 2BGoogle
20.3
–
203 Inflection 3 ProductivityInflection AI
20.1
–
204 GLM-4-32B-0414Z.ai
20.0
–
205 Falcon-H1 1.5B Deep InstructTII
19.9
–
206 Inflection 3 PiInflection AI
19.7
–
207 Qwen3-Omni-30B-A3B-ThinkingAlibaba
19.4
–
208 Qwen2.5-Coder-7B-InstructAlibaba
18.8
–
208 Qwen3-1.7B (reasoning on)Alibaba · best of 2 settings
18.8
–
210 Granite 4.0 H SmallIBM
18.5
–
211 Qwen3-VL-8B-ThinkingAlibaba
18.3
$0.18 / $2.10
212 Qwen3-VL-30B-A3B-ThinkingAlibaba
17.6
$0.29 / $1
213 Qwen3-VL-4B-ThinkingAlibaba
17.4
–
214 Gemma 4 E2BGoogle · best of 2 settings
17.3
–
215 MiniMax-M2MiniMax
17.2
$0.30 / $1.20
216 Solar Pro 3Upstage
16.8
$0.15 / $0.60
217 Llama 2 70B ChatMeta
15.3
–
218 Qwen3-VL-2B-InstructAlibaba
15.2
–
219 Falcon-H1 0.5B InstructTII
14.6
–
220 LFM2-8B-A1BLiquid AI
13.1
–
220 Qwen3-0.6B (reasoning on)Alibaba
13.1
–
222 Qwen3-VL-4B-InstructAlibaba
12.6
–
223 Llama 3.2 1B InstructMeta
11.8
$0.027 / $0.20
223 Qwen3-VL-2B-ThinkingAlibaba
11.8
–
225 Qwen3-4B-Thinking-2507Alibaba
10.9
–
226 Llama 3.2 3B InstructMeta
10.4
$0.05 / $0.33
227 Nanbeige4-3B-Thinking (2511)Nanbeige
9.8
–
227 Llama 3.3 8B InstructAllura Forge
9.8
–
229 Rnj-1 InstructEssential AI
9.0
–
230 Nanbeige4-3B-Thinking (2510)Nanbeige
5.7
–

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Results as published by UGI Leaderboard; we do not re-run them.

What it measures

UGI's blend of intelligence, style, repetition and length adherence in writing, tuned to average human preference.

What it does not measure

Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

947 results from UGI Leaderboard not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Community fine-tunes and merges: 934
  • No score published: 13

UGI Leaderboard by DontPlanToEnd, Apache License 2.0.