Benchmarks / UGI Leaderboard

Reported by UGI Leaderboard

UGI Leaderboard

UGI's blend of intelligence, style, repetition and length adherence in writing, tuned to average human preference.

Last updated 2 Oct 2026

Results dated
6 Sep 2025 to 2 Oct 2026
Results
379 configurations of 230 models
Unit
score out of 100
Licence
Apache License 2.0

Writing score: Claude Opus 5.5

Top 15 of 379 results · score out of 100, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 Gemini 3.8 Flash (medium reasoning)Google 78.6
  2. 2 Gemini 3.7 Flash (medium reasoning)Google 77.5
  3. 3 Gemini 3.7 Flash (high reasoning)Google 76.2
  4. 4 Gemini 3.7 Flash (low reasoning)Google 76.1
  5. 5 Claude Fable 5 (high reasoning)Anthropic 74.2
  6. 6 Claude Fable 5 (max reasoning)Anthropic 73.7
  7. 7 Gemini 3.8 Flash (low reasoning)Google 72.7
  8. 8 Claude Fable 5 (medium reasoning)Anthropic 72.6
  9. 9 Gemini 3.5 Flash (medium reasoning)Google 72.5
  10. 9 Gemini 3.5 Flash (minimal reasoning)Google 72.5
  11. 11 Gemini 3.5 Flash (low reasoning)Google 72.4
  12. 12 Gemini 3.1 Pro Preview (low reasoning)Google 72.2
  13. 13 Gemini 3.8 Flash (high reasoning)Google 72.0
  14. 14 Claude Fable 5 (low reasoning)Anthropic 71.9
  15. 15 Claude Fable 5 (extra-high reasoning)Anthropic 71.7
  16. 22 Claude Opus 5.5 (low reasoning)Anthropic 70.0
  17. 24 Claude Opus 5.5 (medium reasoning)Anthropic 69.8
  18. 28 Claude Opus 5.5 (extra-high reasoning)Anthropic 69.4
  19. 31 Claude Opus 5.5 (max reasoning)Anthropic 68.9
  20. 35 Claude Opus 5.5 (high reasoning)Anthropic 68.4

Full results

UGI: writing score, score out of 100, higher is better
#ModelWriting score
score out of 100, higher is better
Price
$ per million tokens, in / out
1 Gemini 3.8 Flash (medium reasoning)Google
78.6
$1.50 / $7.50
2 Gemini 3.7 Flash (medium reasoning)Google
77.5
$1.50 / $7.50
3 Gemini 3.7 Flash (high reasoning)Google
76.2
$1.50 / $7.50
4 Gemini 3.7 Flash (low reasoning)Google
76.1
$1.50 / $7.50
5 Claude Fable 5 (high reasoning)Anthropic
74.2
$10 / $50
6 Claude Fable 5 (max reasoning)Anthropic
73.7
$10 / $50
7 Gemini 3.8 Flash (low reasoning)Google
72.7
$1.50 / $7.50
8 Claude Fable 5 (medium reasoning)Anthropic
72.6
$10 / $50
9 Gemini 3.5 Flash (medium reasoning)Google
72.5
$1.50 / $9
9 Gemini 3.5 Flash (minimal reasoning)Google
72.5
$1.50 / $9
11 Gemini 3.5 Flash (low reasoning)Google
72.4
$1.50 / $9
12 Gemini 3.1 Pro Preview (low reasoning)Google
72.2
$2 / $12
13 Gemini 3.8 Flash (high reasoning)Google
72.0
$1.50 / $7.50
14 Claude Fable 5 (low reasoning)Anthropic
71.9
$10 / $50
15 Claude Fable 5 (extra-high reasoning)Anthropic
71.7
$10 / $50
16 Claude 3.7 Sonnet (no reasoning)Anthropic
71.2
–
17 Claude Opus 4.6 (high reasoning)Anthropic
70.9
$5 / $25
18 Claude Opus 4.6 (max reasoning)Anthropic
70.7
$5 / $25
19 Claude Opus 4.5 (no reasoning)Anthropic
70.4
$5 / $25
20 Claude Opus 4.5 (reasoning on)Anthropic
70.3
$5 / $25
21 Gemini 3.1 Pro Preview (medium reasoning)Google
70.2
$2 / $12
22 Claude Opus 5.5 (low reasoning)Anthropic
70.0
$4 / $20
23 Claude Opus 4.1 (no reasoning)Anthropic
69.9
$15 / $75
24 Claude Opus 5.5 (medium reasoning)Anthropic
69.8
$4 / $20
24 Gemini 3.6 Flash (medium reasoning)Google
69.8
$1.50 / $7.50
26 Gemini 3.5 Flash (high reasoning)Google
69.7
$1.50 / $9
26 GPT-5.5 (extra-high reasoning)OpenAI
69.7
$5 / $30
28 Claude Opus 5.5 (extra-high reasoning)Anthropic
69.4
$4 / $20
28 Gemini 3.6 Flash (high reasoning)Google
69.4
$1.50 / $7.50
30 Claude Opus 4.6 (medium reasoning)Anthropic
69.0
$5 / $25
31 Claude Opus 5.5 (max reasoning)Anthropic
68.9
$4 / $20
31 GPT-6 Astra (high reasoning)OpenAI
68.9
$10 / $50
33 GPT-5.4 (high reasoning)OpenAI
68.8
$2.50 / $15
34 GPT-5.4 (extra-high reasoning)OpenAI
68.7
$2.50 / $15
35 Claude Opus 5.5 (high reasoning)Anthropic
68.4
$4 / $20
35 DeepSeek-V4-Pro (0423, reasoning on)DeepSeek
68.4
$1.42 / $2.83
37 Claude Opus 4 (no reasoning)Anthropic
68.2
–
37 Gemini 3.1 Pro Preview (high reasoning)Google
68.2
$2 / $12
39 GPT-6 Astra (medium reasoning)OpenAI
68.1
$10 / $50
40 GPT-6 Astra (extra-high reasoning)OpenAI
68.0
$10 / $50
40 Claude 3.7 Sonnet (reasoning on)Anthropic
68.0
–
42 GPT-5.5 (high reasoning)OpenAI
67.9
$5 / $30
43 GPT-6.1 Sol (high reasoning)OpenAI
67.7
$2 / $10
43 Gemini 3 Flash Preview (medium reasoning)Google
67.7
$0.50 / $3
45 GPT-4.1OpenAI
67.5
$2 / $8
46 ChatGPT-4o (2025-03-26)OpenAI
67.4
–
47 GPT-5.4 (medium reasoning)OpenAI
67.3
$2.50 / $15
48 GLM-5.2 (reasoning on)Z.ai
67.0
$1.40 / $4.40
48 Grok 4 (0709)xAI
67.0
–
50 Claude Sonnet 4.5 (no reasoning)Anthropic
66.9
$3 / $15
51 Claude Opus 4 (reasoning on)Anthropic
66.8
–
51 o3 (medium reasoning)OpenAI
66.8
$2 / $8
51 o3 (high reasoning)OpenAI
66.8
$2 / $8
54 Gemini 3 Pro Preview (high reasoning)Google
66.7
–
55 Claude Opus 4.1 (reasoning on)Anthropic
66.5
$15 / $75
56 Claude Opus 4.6 (low reasoning)Anthropic
66.1
$5 / $25
57 Claude Opus 4.8 (max reasoning)Anthropic
65.9
$5 / $25
58 o3 (low reasoning)OpenAI
65.8
$2 / $8
59 Gemini 3.6 Flash (minimal reasoning)Google
65.7
$1.50 / $7.50
60 Claude Opus 4.8 (extra-high reasoning)Anthropic
65.6
$5 / $25
60 Gemini 3 Pro Preview (low reasoning)Google
65.6
–
62 GPT-5.6 Sol (medium reasoning)OpenAI
65.4
$4 / $20
63 GPT-5.6 Sol (high reasoning)OpenAI
65.3
$4 / $20
64 Gemini 2.5 ProGoogle
65.2
$1.25 / $10
65 GPT-6 Astra (low reasoning)OpenAI
65.1
$10 / $50
65 Gemini 3 Flash Preview (minimal reasoning)Google
65.1
$0.50 / $3
67 GPT-6.1 Sol (extra-high reasoning)OpenAI
65.0
$2 / $10
68 GPT-5.4 (low reasoning)OpenAI
64.9
$2.50 / $15
69 GPT-5.5 (medium reasoning)OpenAI
64.8
$5 / $30
70 Claude Opus 4.8 (low reasoning)Anthropic
64.7
$5 / $25
71 Claude Sonnet 4.6 (high reasoning)Anthropic
64.6
$3 / $15
72 Kimi K2.6 (reasoning on)Moonshot AI
64.4
$0.95 / $4
73 Claude Opus 4.8 (high reasoning)Anthropic
64.3
$5 / $25
74 Kimi K2.6 (no reasoning)Moonshot AI
64.2
$0.95 / $4
74 Claude Sonnet 4.5 (reasoning on)Anthropic
64.2
$3 / $15
74 Claude Sonnet 4.6 (max reasoning)Anthropic
64.2
$3 / $15
77 Claude Fable 5.1 (medium reasoning)Anthropic
63.9
$10 / $50
78 Claude Sonnet 4 (no reasoning)Anthropic
63.7
$3 / $15
78 Claude Fable 5.1 (high reasoning)Anthropic
63.7
$10 / $50
80 Claude Opus 4.8 (medium reasoning)Anthropic
63.5
$5 / $25
81 GPT-4o (2024-05-13)OpenAI
63.4
$5 / $15
82 Grok 4.6xAI
63.3
$2 / $6
82 Gemini 3 Flash Preview (high reasoning)Google
63.3
$0.50 / $3
84 GPT-5.6 Sol (extra-high reasoning)OpenAI
63.2
$4 / $20
85 Grok 4.20 Multi-Agent Beta (0309, 4 agents)xAI
63.1
–
86 GPT-5.6 Sol (low reasoning)OpenAI
63.0
$4 / $20
87 Claude Sonnet 5 (max reasoning)Anthropic
62.7
$2 / $10
87 GPT-5 Chat (latest)OpenAI
62.7
–
89 GPT-6.1 Sol (medium reasoning)OpenAI
62.4
$2 / $10
89 Kimi K2.5 (reasoning on)Moonshot AI
62.4
$0.57 / $2.85
91 Claude Sonnet 4 (reasoning on)Anthropic
62.3
$3 / $15
92 GLM-5.1 (reasoning on)Z.ai
61.7
$1.38 / $4.40
93 Kimi K2.5 (no reasoning)Moonshot AI
61.5
$0.57 / $2.85
94 Claude Opus 4.7 (max reasoning)Anthropic
61.4
$5 / $25
95 Claude Sonnet 4.6 (medium reasoning)Anthropic
61.2
$3 / $15
95 GLM-5.2 (no reasoning)Z.ai
61.2
$1.40 / $4.40
97 o1 (low reasoning)OpenAI
61.0
$15 / $60
98 GPT-5.2 ChatOpenAI
60.8
–
99 GPT-5.3 ChatOpenAI
60.7
–
100 GPT-6.1 Sol (low reasoning)OpenAI
60.5
$2 / $10
101 GPT-5.6 Terra (high reasoning)OpenAI
60.4
$2 / $12
102 DeepSeek-V4-Pro (0423, no reasoning)DeepSeek
60.3
$1.42 / $2.83
102 GPT-5.1 ChatOpenAI
60.3
–
104 Claude Fable 5.1 (low reasoning)Anthropic
60.2
$10 / $50
105 o1 (high reasoning)OpenAI
60.1
$15 / $60
106 Grok 4.5xAI
60.0
$2 / $6
107 GPT-5.6 Sol (no reasoning)OpenAI
59.9
$4 / $20
108 Claude Sonnet 4.6 (low reasoning)Anthropic
59.3
$3 / $15
108 Kimi K2 ThinkingMoonshot AI
59.3
$0.60 / $2.50
110 GPT-5.6 Terra (extra-high reasoning)OpenAI
58.9
$2 / $12
111 GPT-5.6 Terra (medium reasoning)OpenAI
58.8
$2 / $12
112 Grok 3xAI
58.7
–
112 GPT-5.5 (low reasoning)OpenAI
58.7
$5 / $30
114 Claude Sonnet 5 (extra-high reasoning)Anthropic
58.5
$2 / $10
115 Qwen3.6-Plus (reasoning on)Alibaba
58.2
$0.33 / $1.95
116 GPT-5.4 (no reasoning)OpenAI
58.0
$2.50 / $15
117 Grok 4.3xAI
57.7
$1.25 / $2.50
118 MiMo-V2.5-Pro (reasoning on)Xiaomi
57.4
$0.43 / $0.87
118 Grok 4.20 Beta (0309, reasoning)xAI
57.4
–
120 GLM-4.5 (reasoning on)Z.ai
57.3
$0.60 / $2.20
120 DeepSeek-V3.2 (reasoning on)DeepSeek
57.3
$0.30 / $0.96
122 Claude Opus 4.7 (low reasoning)Anthropic
56.9
$5 / $25
122 Claude Opus 4.7 (high reasoning)Anthropic
56.9
$5 / $25
124 GPT-5.6 Terra (low reasoning)OpenAI
56.8
$2 / $12
124 Claude Sonnet 5 (high reasoning)Anthropic
56.8
$2 / $10
126 Claude Opus 4.7 (medium reasoning)Anthropic
56.6
$5 / $25
127 MiMo-V2.5 (reasoning on)Xiaomi
56.4
$0.17 / $0.34
128 DeepSeek-V3.2 (no reasoning)DeepSeek
56.3
$0.30 / $0.96
129 GPT-5.1 (high reasoning)OpenAI
55.8
$1.25 / $10
130 GPT-6 Sol (high reasoning)OpenAI
55.6
$2 / $10
131 Grok 4.20 (0309, reasoning)xAI
55.3
–
132 GLM-5 (reasoning on)Z.ai
55.0
$0.95 / $2.55
133 GLM-5.1 (no reasoning)Z.ai
54.9
$1.38 / $4.40
134 DeepSeek-V4-Flash (0423, no reasoning)DeepSeek
54.6
$0.14 / $0.28
135 Kimi K2 (0905)Moonshot AI
54.4
$0.60 / $2.50
136 Claude 3 OpusAnthropic
54.3
–
137 GLM-4.5 (no reasoning)Z.ai
54.2
$0.60 / $2.20
138 Grok 4.7xAI
54.1
$2 / $6
139 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek
54.0
$0.27 / $1
139 DeepSeek-V3.2-Exp (reasoning on)DeepSeek
54.0
$0.27 / $0.41
139 DeepSeek-V3.2-SpecialeDeepSeek
54.0
–
142 GPT-5.1 (medium reasoning)OpenAI
53.9
$1.25 / $10
142 GPT-6 Sol (extra-high reasoning)OpenAI
53.9
$2 / $10
144 GPT-6 Sol (low reasoning)OpenAI
53.8
$2 / $10
145 GPT-5.6 Luna (high reasoning)OpenAI
53.7
$0.20 / $1.20
146 GPT-6 Sol (medium reasoning)OpenAI
53.6
$2 / $10
147 GPT-5 (high reasoning)OpenAI
53.5
$1.25 / $10
148 Grok 4 Fast (reasoning)xAI
53.4
–
149 GLM-4.6 (reasoning on)Z.ai
53.3
$0.50 / $2
150 Claude Sonnet 5 (low reasoning)Anthropic
53.0
$2 / $10
151 GPT-5.6 Luna (extra-high reasoning)OpenAI
52.8
$0.20 / $1.20
152 Kimi K2 (0711)Moonshot AI
52.5
$0.57 / $2.30
153 Claude Sonnet 5 (medium reasoning)Anthropic
52.4
$2 / $10
154 Claude Haiku 4.5 (no reasoning)Anthropic
52.2
$1 / $5
155 DeepSeek-V3.1-Terminus (no reasoning)DeepSeek
52.0
$0.27 / $1
156 DeepSeek-V3.1 (reasoning on)DeepSeek
51.9
$0.55 / $1.65
157 DeepSeek-R1DeepSeek
51.8
$0.70 / $2.50
158 MiMo-V2.5-Pro (no reasoning)Xiaomi
51.7
$0.43 / $0.87
159 GPT-5.5 (no reasoning)OpenAI
51.5
$5 / $30
160 Grok 4.1 Fast (reasoning)xAI
51.3
–
161 Hy3 Preview (reasoning on)Tencent
51.1
$0.18 / $0.60
161 DeepSeek-V3.1 (no reasoning)DeepSeek
51.1
$0.55 / $1.65
163 DeepSeek-V3.2-Exp (no reasoning)DeepSeek
51.0
$0.27 / $0.41
164 GPT-5.1 (low reasoning)OpenAI
50.8
$1.25 / $10
165 Qwen3.5-397B-A17B (no reasoning)Alibaba
50.1
$0.55 / $3.50
166 GLM-4.6 (no reasoning)Z.ai
50.0
$0.50 / $2
167 DeepSeek-V4-Flash (0423, reasoning on)DeepSeek
49.7
$0.14 / $0.28
168 GPT-5.2 (high reasoning)OpenAI
49.4
$1.75 / $14
169 GPT-5.6 Luna (medium reasoning)OpenAI
49.2
$0.20 / $1.20
169 GLM-5 (no reasoning)Z.ai
49.2
$0.95 / $2.55
171 Claude Haiku 4.5 (reasoning on)Anthropic
49.1
$1 / $5
171 DeepSeek-V3-0324DeepSeek
49.1
$0.25 / $1
173 Qwen3-235B-A22B-Instruct-2507Alibaba
49.0
$0.15 / $0.75
174 GPT-5 (low reasoning)OpenAI
48.3
$1.25 / $10
175 GPT-5.6 Terra (no reasoning)OpenAI
48.2
$2 / $12
176 GLM-4.7 (reasoning on)Z.ai
47.4
$0.54 / $1.98
177 Grok 4.20 (0309, non-reasoning)xAI
47.2
–
178 GLM-4.5-Air (no reasoning)Z.ai
47.0
$0.14 / $0.86
178 GPT-5.6 Luna (low reasoning)OpenAI
47.0
$0.20 / $1.20
180 o4-mini (low reasoning)OpenAI
46.9
$1.10 / $4.40
181 Qwen3.6-Plus (no reasoning)Alibaba
46.1
$0.33 / $1.95
182 Step 3.5 FlashStepFun
45.9
$0.10 / $0.30
183 Mistral Medium 3.5 (no reasoning)Mistral AI
45.5
$1.50 / $7.50
184 o4-mini (medium reasoning)OpenAI
45.4
$1.10 / $4.40
184 Qwen3.6-35B-A3B (reasoning by prefill)Alibaba
45.4
$0.10 / $1
186 Gemma 3 27BGoogle
45.0
$0.12 / $0.20
187 Gemini 2.5 Flash Preview (09-2025, no reasoning)Google
44.9
–
187 MiMo-V2-Flash (reasoning on)Xiaomi
44.9
–
189 MiMo-V2.5 (no reasoning)Xiaomi
44.8
$0.17 / $0.34
190 GLM-4.7 (no reasoning)Z.ai
44.2
$0.54 / $1.98
191 GPT-5.2 (medium reasoning)OpenAI
44.1
$1.75 / $14
192 Gemini 2.5 Flash Preview (09-2025, reasoning on)Google
43.9
–
192 Grok 4.20 Beta (0309, non-reasoning)xAI
43.9
–
192 Mistral Medium 3.5 (high reasoning)Mistral AI
43.9
$1.50 / $7.50
192 GPT-5.2 (low reasoning)OpenAI
43.9
$1.75 / $14
196 Gemma 4 26B A4B (reasoning by prefill)Google
43.8
$0.10 / $0.30
197 DeepSeek-V3DeepSeek
43.2
$0.26 / $1.03
197 o4-mini (high reasoning)OpenAI
43.2
$1.10 / $4.40
199 Gemma 4 31B (reasoning by prefill)Google
43.1
$0.14 / $0.40
200 Grok 4 Fast (non-reasoning)xAI
43.0
–
200 Qwen3-VL-235B-A22B-InstructAlibaba
43.0
$0.30 / $1.50
202 Qwen3.5-35B-A3B (reasoning by prefill)Alibaba
42.6
$0.16 / $1.30
203 Qwen3.6-27B (reasoning by prefill)Alibaba
42.5
$0.30 / $3.20
204 Qwen3.5-27B (reasoning by prefill)Alibaba
42.4
$0.27 / $2.16
205 DeepSeek-R1-0528DeepSeek
42.3
$0.50 / $2.18
206 GLM-4.5-AirZ.ai
42.0
$0.14 / $0.86
207 Grok 4.1 Fast (non-reasoning)xAI
41.9
–
208 Qwen3-Next-80B-A3B-InstructAlibaba
41.8
$0.10 / $1.10
209 MiniMax-M2.5MiniMax
41.7
$0.30 / $1.20
210 MiniMax-M2.1MiniMax
41.6
$0.30 / $1.20
210 Mistral Large 3 (no reasoning)Mistral AI
41.6
$0.50 / $1.50
210 MiMo-V2-Flash (no reasoning)Xiaomi
41.6
–
210 Gemma 4 26B A4BGoogle
41.6
$0.10 / $0.30
210 Hy3 Preview (no reasoning)Tencent
41.6
$0.18 / $0.60
215 Mistral Large 3 (reasoning on)Mistral AI
41.5
$0.50 / $1.50
216 Muse Glimmer 30B (extra-high reasoning)Meta
41.0
$0.30 / $1.20
216 Muse Glimmer 30B (high reasoning)Meta
41.0
$0.30 / $1.20
216 GPT-5.6 Luna (no reasoning)OpenAI
41.0
$0.20 / $1.20
219 Mistral Large 2 (2407)Mistral AI
40.9
$2 / $6
220 Trinity Large PreviewArcee AI
40.7
–
221 Command ACohere
40.6
$2.50 / $10
222 Mistral Small 4 (no reasoning)Mistral AI
40.3
$0.15 / $0.60
223 Muse Glimmer 30B (medium reasoning)Meta
40.0
$0.30 / $1.20
224 Qwen3.5-27B (no reasoning)Alibaba
39.9
$0.27 / $2.16
225 GLM-4.6V (reasoning on)Z.ai
39.7
$0.30 / $0.90
226 Qwen3.5-122B-A10B (no reasoning)Alibaba
39.5
$0.26 / $2.08
226 Qwen3.5-9B (reasoning by prefill)Alibaba
39.5
$0.10 / $0.15
226 Mistral Medium 3.1Mistral AI
39.5
$0.40 / $2
229 GPT-5.2 (no reasoning)OpenAI
39.2
$1.75 / $14
230 Mistral Small 4 (high reasoning)Mistral AI
39.0
$0.15 / $0.60
231 Qwen3.6-27B (no reasoning)Alibaba
38.8
$0.30 / $3.20
232 Qwen3.5-122B-A10B (reasoning on)Alibaba
38.7
$0.26 / $2.08
233 Qwen3-MaxAlibaba
38.6
$0.78 / $3.90
233 Gemma 4 31BGoogle
38.6
$0.14 / $0.40
235 gpt-oss-120b (medium reasoning)OpenAI
38.5
$0.15 / $0.60
236 Jamba Large 1.7AI21 Labs
38.1
–
237 Mistral Large 2.1 (2411)Mistral AI
37.8
–
238 Kimi Linear 48B A3B InstructMoonshot AI
37.7
–
239 MiniMax-M2.7MiniMax
37.6
$0.30 / $1.20
240 Qwen3.5-35B-A3B (no reasoning)Alibaba
37.0
$0.16 / $1.30
241 Gemma 4 12B (reasoning by prefill)Google
36.3
–
241 Mistral Small 3.2 24BMistral AI
36.3
$0.094 / $0.25
243 MedGemma 27B TextGoogle
36.1
–
244 Qwen3.6-35B-A3B (no reasoning)Alibaba
35.8
$0.10 / $1
245 Qwen2.5-VL-72B-InstructAlibaba
35.6
$0.80 / $1
246 GLM-4.6V (no reasoning)Z.ai
35.2
$0.30 / $0.90
246 Muse Glimmer 30B (low reasoning)Meta
35.2
$0.30 / $1.20
248 Gemma 2 27BGoogle
35.0
$0.65 / $0.65
248 Qwen3-VL-32B-InstructAlibaba
35.0
$0.10 / $0.42
250 QwQ-32BAlibaba
34.9
–
251 Qwen3-14B (no reasoning)Alibaba
34.8
$0.12 / $0.24
252 Qwen2.5-32B-InstructAlibaba
34.2
–
253 Magistral Small 1.2Mistral AI
34.0
–
254 Mistral Small 3Mistral AI
33.9
$0.05 / $0.08
255 Qwen3.5-9B (no reasoning)Alibaba
33.5
$0.10 / $0.15
256 Nemotron Nano 12B v2 VL (BF16)NVIDIA
33.3
–
257 Mistral NemoMistral AI
33.1
$0.023 / $0.03
258 Qwen3-32B (no reasoning)Alibaba
33.0
$0.14 / $0.40
259 Llama 3.1 70B InstructMeta
32.8
$0.40 / $0.40
260 Mistral Medium 3Mistral AI
32.6
$0.40 / $2
260 Mistral Small (2409)Mistral AI
32.6
–
262 Ring-1TAnt Group
32.4
–
262 Qwen3-235B-A22B-Thinking-2507Alibaba
32.4
$0.30 / $3
264 Qwen3-VL-235B-A22B-ThinkingAlibaba
32.3
$0.40 / $4
264 Command R+ (08-2024)Cohere
32.3
$2.50 / $10
266 Mistral Small 3.1 24BMistral AI
32.1
$0.35 / $0.56
267 Seed-OSS-36B-Instruct (512-token reasoning budget)ByteDance
32.0
–
268 Claude 3 HaikuAnthropic
31.9
–
269 Qwen3-Coder-30B-A3B-InstructAlibaba
31.8
$0.07 / $0.28
269 Nemotron Nano 12B v2 VL (BF16, no reasoning)NVIDIA
31.8
–
271 Gemma 4 12BGoogle
31.6
–
271 Devstral Small 2Mistral AI
31.6
–
273 Qwen2.5-72B-InstructAlibaba
31.5
$0.36 / $0.40
274 Ling-1TAnt Group
31.4
–
274 Ministral 3 14BMistral AI
31.4
$0.20 / $0.20
276 Qwen3.5-4B (reasoning by prefill)Alibaba
30.8
–
277 Qwen3-32B (reasoning on)Alibaba
30.3
$0.14 / $0.40
277 Qwen3-VL-8B-InstructAlibaba
30.3
$0.12 / $0.46
279 Qwen3-30B-A3B (no reasoning)Alibaba
30.2
$0.12 / $0.50
280 Apriel Nemotron 15B ThinkerServiceNow
30.0
–
281 Qwen3-4B-Instruct-2507Alibaba
29.9
–
281 Llama 3.1 405B InstructMeta
29.9
–
281 Gemma 3 12BGoogle
29.9
$0.05 / $0.15
284 Qwen2.5-14B-InstructAlibaba
29.8
–
285 Qwen2.5-7B-InstructAlibaba
29.7
$0.10 / $0.20
285 Qwen3.5-4B (no reasoning)Alibaba
29.7
–
287 Qwen3-14B (reasoning on)Alibaba
29.6
$0.12 / $0.24
288 Qwen3.5-2B (no reasoning)Alibaba
29.0
–
289 Ministral 3 8BMistral AI
28.8
$0.15 / $0.15
290 Qwen3-30B-A3B (reasoning on)Alibaba
28.5
$0.12 / $0.50
290 Olmo 3 32B ThinkAi2
28.5
–
292 Seed-OSS-36B-Instruct (no reasoning)ByteDance
28.2
–
293 Qwen3-8B (no reasoning)Alibaba
28.0
$0.12 / $0.46
294 Qwen3-30B-A3B-Instruct-2507Alibaba
27.9
$0.09 / $0.30
295 Ministral 3 8B (reasoning by prefill)Mistral AI
27.6
$0.15 / $0.15
295 LFM2-24B-A2BLiquid AI
27.6
–
297 Seed-OSS-36B-InstructByteDance
27.5
–
298 Mixtral 8x22B InstructMistral AI
27.4
$2 / $6
299 Ministral 3 8B ReasoningMistral AI
27.2
–
299 Qwen3-Next-80B-A3B-ThinkingAlibaba
27.2
$0.15 / $1.20
301 InternLM3 8B InstructShanghai AI Lab
27.1
–
302 Llama 4 ScoutMeta
26.7
$0.18 / $0.59
303 Command R (08-2024)Cohere
26.5
$0.15 / $0.60
304 Llama 3.3 70B InstructMeta
26.2
$0.59 / $0.79
305 Olmo 3 7B ThinkAi2
26.0
–
306 Nemotron 3 Nano 30B A3BNVIDIA
25.9
$0.05 / $0.20
307 Ministral 3 14B (reasoning by prefill)Mistral AI
25.7
$0.20 / $0.20
307 Falcon-H1 7B InstructTII
25.7
–
307 Phi-4Microsoft
25.7
$0.07 / $0.14
307 Mixtral 8x7B Instruct v0.1Mistral AI
25.7
–
311 Ministral 3 14B ReasoningMistral AI
25.6
–
311 EXAONE 4.0 32BLG AI Research
25.6
–
313 GLM-4.7-Flash (no reasoning)Z.ai
25.0
$0.06 / $0.40
313 Gemma 2 9BGoogle
25.0
–
315 Llama 3.1 8B InstructMeta
24.8
$0.05 / $0.08
316 gpt-oss-20b (low reasoning)OpenAI
24.6
$0.03 / $0.15
317 gpt-oss-20b (medium reasoning)OpenAI
24.5
$0.03 / $0.15
318 Kimi-VL-A3B-InstructMoonshot AI
24.4
–
318 Olmo 3 7B InstructAi2
24.4
–
320 Gemma 3 4BGoogle
24.0
$0.05 / $0.10
320 Llama 4 MaverickMeta
24.0
$0.27 / $0.85
322 Qwen3-8B (reasoning on)Alibaba
23.9
$0.12 / $0.46
323 Jamba Mini 1.7AI21 Labs
23.8
–
323 Reka Flash 3Reka AI
23.8
$0.10 / $0.20
325 Solar 10.7B Instruct v1.0Upstage
23.6
–
326 Nova 2 Lite (no reasoning)Amazon
23.5
$0.30 / $2.50
327 Nova 2 Lite (reasoning on)Amazon
23.4
$0.30 / $2.50
328 Qwen3.5-0.8B (no reasoning)Alibaba
23.0
–
329 Qwen3-VL-32B-ThinkingAlibaba
22.9
–
330 Falcon-H1 3B InstructTII
22.8
–
331 Qwen3-4B (no reasoning)Alibaba
22.4
–
332 Qwen2.5-1.5B-InstructAlibaba
22.2
–
333 Qwen2.5-VL-32B-InstructAlibaba
21.7
–
334 Gemma 4 E4B (reasoning by prefill)Google
21.6
–
335 Falcon-H1 1.5B InstructTII
21.0
–
336 Qwen3-30B-A3B-Thinking-2507Alibaba
20.6
$0.20 / $2.40
336 Kimi-VL-A3B-Thinking (2506)Moonshot AI
20.6
–
338 Qwen2.5-VL-3B-InstructAlibaba
20.4
–
339 Gemma 2 2BGoogle
20.3
–
340 Gemma 4 E4BGoogle
20.2
–
341 Inflection 3 ProductivityInflection AI
20.1
–
342 GLM-4-32B-0414Z.ai
20.0
–
343 Falcon-H1 1.5B Deep InstructTII
19.9
–
344 Inflection 3 PiInflection AI
19.7
–
345 Qwen3-Omni-30B-A3B-ThinkingAlibaba
19.4
–
346 Qwen3.5-2B (reasoning by prefill)Alibaba
19.3
–
347 Ministral 3 14B Reasoning (reasoning by prefill)Mistral AI
19.2
–
347 Magistral Small 1.2 (reasoning by system prompt)Mistral AI
19.2
–
349 Qwen3-4B (reasoning on)Alibaba
18.8
–
349 Qwen2.5-Coder-7B-InstructAlibaba
18.8
–
349 Qwen3-1.7B (reasoning on)Alibaba
18.8
–
352 Granite 4.0 H SmallIBM
18.5
–
353 Qwen3-VL-8B-ThinkingAlibaba
18.3
$0.18 / $2.10
354 Ministral 3 8B Reasoning (reasoning by prefill)Mistral AI
17.8
–
355 Qwen3-VL-30B-A3B-ThinkingAlibaba
17.6
$0.29 / $1
356 Qwen3-VL-4B-ThinkingAlibaba
17.4
–
357 Gemma 4 E2BGoogle
17.3
–
358 MiniMax-M2MiniMax
17.2
$0.30 / $1.20
359 Solar Pro 3Upstage
16.8
$0.15 / $0.60
360 Gemma 4 E2B (reasoning by prefill)Google
16.2
–
361 Llama 2 70B ChatMeta
15.3
–
362 Qwen3-VL-2B-InstructAlibaba
15.2
–
363 GLM-4.7-Flash (reasoning by prefill)Z.ai
14.7
$0.06 / $0.40
364 Falcon-H1 0.5B InstructTII
14.6
–
365 LFM2-8B-A1BLiquid AI
13.1
–
365 Qwen3-0.6B (reasoning on)Alibaba
13.1
–
367 Qwen3-VL-4B-InstructAlibaba
12.6
–
368 Qwen3-1.7B (no reasoning)Alibaba
12.2
–
369 Llama 3.2 1B InstructMeta
11.8
$0.027 / $0.20
369 Qwen3-VL-2B-ThinkingAlibaba
11.8
–
371 gpt-oss-20b (high reasoning)OpenAI
10.9
$0.03 / $0.15
371 Qwen3-4B-Thinking-2507Alibaba
10.9
–
373 Llama 3.2 3B InstructMeta
10.4
$0.05 / $0.33
374 Nanbeige4-3B-Thinking (2511)Nanbeige
9.8
–
374 Llama 3.3 8B InstructAllura Forge
9.8
–
376 Rnj-1 InstructEssential AI
9.0
–
377 Ministral 3 8B Reasoning (reasoning by system prompt)Mistral AI
6.1
–
378 Nanbeige4-3B-Thinking (2510)Nanbeige
5.7
–
379 Nemotron 3 Nano 30B A3B (reasoning by prefill)NVIDIA
3.4
$0.05 / $0.20

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Results as published by UGI Leaderboard; we do not re-run them.

What it measures

UGI's blend of intelligence, style, repetition and length adherence in writing, tuned to average human preference.

What it does not measure

Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

947 results from UGI Leaderboard not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Community fine-tunes and merges: 934
  • No score published: 13

UGI Leaderboard by DontPlanToEnd, Apache License 2.0.