Benchmarks / UGI Leaderboard

Reported by UGI Leaderboard

UGI Leaderboard

Average percentage difference between a requested and the delivered word count.

Last updated 2 Oct 2026

Results dated
6 Sep 2025 to 2 Oct 2026
Results
379 configurations of 230 models
Unit
% off the requested word count
Licence
Apache License 2.0

Requested-length error: Claude Opus 5.5

Top 15 of 230 results · % off the requested word count, lower is better · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ GPT-6.1 Sol (high reasoning)OpenAI 0.0%
  2. 1≈ GPT-6 Astra (high reasoning)OpenAI 0.0%
  3. 3 Claude Opus 4.7 (low reasoning)Anthropic 1.0%
  4. 3 GPT-5.6 Sol (extra-high reasoning)OpenAI 1.0%
  5. 5 Claude Fable 5 (high reasoning)Anthropic 2.0%
  6. 5 Claude Opus 5.5 (max reasoning)Anthropic 2.0%
  7. 5 Claude Sonnet 4.6 (low reasoning)Anthropic 2.0%
  8. 8 Claude Fable 5.1 (high reasoning)Anthropic 3.0%
  9. 8 Claude Opus 4.1 (no reasoning)Anthropic 3.0%
  10. 8 Claude Opus 4.5 (no reasoning)Anthropic 3.0%
  11. 8 Claude Opus 4.6 (max reasoning)Anthropic 3.0%
  12. 8 Claude Sonnet 5 (high reasoning)Anthropic 3.0%
  13. 13 Claude Sonnet 4 (no reasoning)Anthropic 4.0%
  14. 13 GPT-6 Sol (high reasoning)OpenAI 4.0%
  15. 15 Claude Sonnet 4.5 (no reasoning)Anthropic 5.0%

Full results

UGI: requested-length error, % off the requested word count, lower is better
#ModelRequested-length error
% off the requested word count, lower is better
Price
$ per million tokens, in / out
1≈ GPT-6.1 Sol (high reasoning)OpenAI · best of 4 settings
0.0%
$2 / $10
1≈ GPT-6 Astra (high reasoning)OpenAI · best of 4 settings
0.0%
$10 / $50
3 Claude Opus 4.7 (low reasoning)Anthropic · best of 4 settings
1.0%
$5 / $25
3 GPT-5.6 Sol (extra-high reasoning)OpenAI · best of 5 settings
1.0%
$4 / $20
5 Claude Fable 5 (high reasoning)Anthropic · best of 5 settings
2.0%
$10 / $50
5 Claude Opus 5.5 (max reasoning)Anthropic · best of 5 settings
2.0%
$4 / $20
5 Claude Sonnet 4.6 (low reasoning)Anthropic · best of 4 settings
2.0%
$3 / $15
8 Claude Fable 5.1 (high reasoning)Anthropic · best of 3 settings
3.0%
$10 / $50
8 Claude Opus 4.1 (no reasoning)Anthropic · best of 2 settings
3.0%
$15 / $75
8 Claude Opus 4.5 (no reasoning)Anthropic · best of 2 settings
3.0%
$5 / $25
8 Claude Opus 4.6 (max reasoning)Anthropic · best of 4 settings
3.0%
$5 / $25
8 Claude Sonnet 5 (high reasoning)Anthropic · best of 5 settings
3.0%
$2 / $10
13 Claude Sonnet 4 (no reasoning)Anthropic · best of 2 settings
4.0%
$3 / $15
13 GPT-6 Sol (high reasoning)OpenAI · best of 4 settings
4.0%
$2 / $10
15 Claude Sonnet 4.5 (no reasoning)Anthropic · best of 2 settings
5.0%
$3 / $15
15 Gemma 4 31BGoogle · best of 2 settings
5.0%
$0.14 / $0.40
17 Claude Haiku 4.5 (no reasoning)Anthropic · best of 2 settings
6.0%
$1 / $5
17 Claude Opus 4.8 (extra-high reasoning)Anthropic · best of 5 settings
6.0%
$5 / $25
17 Claude Opus 4 (no reasoning)Anthropic · best of 2 settings
6.0%
–
17 GPT-4.1OpenAI
6.0%
$2 / $8
21 Qwen3-Next-80B-A3B-InstructAlibaba
7.0%
$0.10 / $1.10
21 Claude 3.7 Sonnet (no reasoning)Anthropic · best of 2 settings
7.0%
–
21 Gemini 3.7 Flash (medium reasoning)Google · best of 3 settings
7.0%
$1.50 / $7.50
21 Grok 4.1 Fast (reasoning)xAI
7.0%
–
25 Qwen3-4B-Instruct-2507Alibaba
8.0%
–
25 Qwen3.6-Plus (reasoning on)Alibaba · best of 2 settings
8.0%
$0.33 / $1.95
25 Gemma 3 27BGoogle
8.0%
$0.12 / $0.20
25 Gemma 4 12B (reasoning by prefill)Google · best of 2 settings
8.0%
–
25 Gemma 4 26B A4B (reasoning by prefill)Google · best of 2 settings
8.0%
$0.10 / $0.30
25 Mistral Small 3Mistral AI
8.0%
$0.05 / $0.08
25 GPT-5.3 ChatOpenAI
8.0%
–
25 MiMo-V2.5-Pro (reasoning on)Xiaomi · best of 2 settings
8.0%
$0.43 / $0.87
33 Qwen3-4B (no reasoning)Alibaba · best of 2 settings
9.0%
–
33 Qwen3.5-122B-A10B (reasoning on)Alibaba · best of 2 settings
9.0%
$0.26 / $2.08
33 Qwen3.6-27B (no reasoning)Alibaba · best of 2 settings
9.0%
$0.30 / $3.20
33 Gemini 3.5 Flash (low reasoning)Google · best of 4 settings
9.0%
$1.50 / $9
33 Gemini 3.6 Flash (minimal reasoning)Google · best of 3 settings
9.0%
$1.50 / $7.50
33 Gemma 3 12BGoogle
9.0%
$0.05 / $0.15
33 Gemma 3 4BGoogle
9.0%
$0.05 / $0.10
33 Gemma 4 E2BGoogle · best of 2 settings
9.0%
–
33 ChatGPT-4o (2025-03-26)OpenAI
9.0%
–
33 GLM-5.1 (reasoning on)Z.ai · best of 2 settings
9.0%
$1.38 / $4.40
43 Qwen3.6-35B-A3B (no reasoning)Alibaba · best of 2 settings
10.0%
$0.10 / $1
43 Gemma 4 E4B (reasoning by prefill)Google · best of 2 settings
10.0%
–
43 Llama 3.3 70B InstructMeta
10.0%
$0.59 / $0.79
43 Kimi K2.5 (reasoning on)Moonshot AI · best of 2 settings
10.0%
$0.57 / $2.85
43 Grok 4.1 Fast (non-reasoning)xAI
10.0%
–
43 GLM-4.5-Air (no reasoning)Z.ai · best of 2 settings
10.0%
$0.14 / $0.86
49 Qwen3.5-27B (reasoning by prefill)Alibaba · best of 2 settings
11.0%
$0.27 / $2.16
49 Mistral Large 2 (2407)Mistral AI
11.0%
$2 / $6
49 Mistral Large 2.1 (2411)Mistral AI
11.0%
–
49 Mistral Small 3.1 24BMistral AI
11.0%
$0.35 / $0.56
49 Mistral Small (2409)Mistral AI
11.0%
–
54 DeepSeek-V4-Flash (0423, reasoning on)DeepSeek · best of 2 settings
12.0%
$0.14 / $0.28
54 GPT-4o (2024-05-13)OpenAI
12.0%
$5 / $15
54 GLM-5.2 (no reasoning)Z.ai · best of 2 settings
12.0%
$1.40 / $4.40
57 Qwen3-30B-A3B-Instruct-2507Alibaba
13.0%
$0.09 / $0.30
57 Gemini 3.8 Flash (high reasoning)Google · best of 3 settings
13.0%
$1.50 / $7.50
57 Llama 4 MaverickMeta
13.0%
$0.27 / $0.85
57 GPT-5.5 (extra-high reasoning)OpenAI · best of 5 settings
13.0%
$5 / $30
57 GLM-4.5 (no reasoning)Z.ai · best of 2 settings
13.0%
$0.60 / $2.20
57 GLM-4.6 (reasoning on)Z.ai · best of 2 settings
13.0%
$0.50 / $2
63 Qwen3-30B-A3B (no reasoning)Alibaba · best of 2 settings
14.0%
$0.12 / $0.50
63 MedGemma 27B TextGoogle
14.0%
–
63 LFM2-24B-A2BLiquid AI
14.0%
–
63 Llama 3.1 405B InstructMeta
14.0%
–
63 Muse Glimmer 30B (extra-high reasoning)Meta · best of 4 settings
14.0%
$0.30 / $1.20
63 GPT-5.1 ChatOpenAI
14.0%
–
69 Qwen3-235B-A22B-Instruct-2507Alibaba
15.0%
$0.15 / $0.75
69 Qwen3.5-9B (reasoning by prefill)Alibaba · best of 2 settings
15.0%
$0.10 / $0.15
69 Olmo 3 7B InstructAi2
15.0%
–
69 Claude 3 HaikuAnthropic
15.0%
–
69 Llama 3.1 70B InstructMeta
15.0%
$0.40 / $0.40
69 Mistral Large 3 (no reasoning)Mistral AI · best of 2 settings
15.0%
$0.50 / $1.50
69 GPT-5.6 Luna (extra-high reasoning)OpenAI · best of 5 settings
15.0%
$0.20 / $1.20
76 Qwen3.5-2B (no reasoning)Alibaba · best of 2 settings
16.0%
–
76 Qwen3.5-4B (reasoning by prefill)Alibaba · best of 2 settings
16.0%
–
76 Qwen3-Coder-30B-A3B-InstructAlibaba
16.0%
$0.07 / $0.28
76 Gemini 3 Flash Preview (minimal reasoning)Google · best of 3 settings
16.0%
$0.50 / $3
76 GPT-5 Chat (latest)OpenAI
16.0%
–
76 Grok 3xAI
16.0%
–
76 MiMo-V2.5 (reasoning on)Xiaomi · best of 2 settings
16.0%
$0.17 / $0.34
83 Qwen2.5-32B-InstructAlibaba
17.0%
–
83 Qwen3.5-35B-A3B (no reasoning)Alibaba · best of 2 settings
17.0%
$0.16 / $1.30
83 Qwen3-8B (no reasoning)Alibaba · best of 2 settings
17.0%
$0.12 / $0.46
83 DeepSeek-V3DeepSeek
17.0%
$0.26 / $1.03
83 Magistral Small 1.2Mistral AI · best of 2 settings
17.0%
–
83 Mistral Small 3.2 24BMistral AI
17.0%
$0.094 / $0.25
83 Falcon-H1 7B InstructTII
17.0%
–
83 Grok 4 (0709)xAI
17.0%
–
91 Qwen3-0.6B (reasoning on)Alibaba
18.0%
–
91 DeepSeek-R1DeepSeek
18.0%
$0.70 / $2.50
91 Mistral Medium 3.5 (no reasoning)Mistral AI · best of 2 settings
18.0%
$1.50 / $7.50
91 Mistral NemoMistral AI
18.0%
$0.023 / $0.03
91 Kimi K2.6 (reasoning on)Moonshot AI · best of 2 settings
18.0%
$0.95 / $4
91 o3 (high reasoning)OpenAI · best of 3 settings
18.0%
$2 / $8
91 Hy3 Preview (reasoning on)Tencent · best of 2 settings
18.0%
$0.18 / $0.60
91 Falcon-H1 3B InstructTII
18.0%
–
99 Qwen2.5-7B-InstructAlibaba
19.0%
$0.10 / $0.20
99 Qwen2.5-14B-InstructAlibaba
19.0%
–
99 Qwen3-32B (no reasoning)Alibaba · best of 2 settings
19.0%
$0.14 / $0.40
99 Qwen3-VL-235B-A22B-InstructAlibaba
19.0%
$0.30 / $1.50
99 Qwen3-VL-2B-InstructAlibaba
19.0%
–
99 Gemini 2.5 ProGoogle
19.0%
$1.25 / $10
99 Gemini 3.1 Pro Preview (low reasoning)Google · best of 3 settings
19.0%
$2 / $12
99 Gemma 2 2BGoogle
19.0%
–
99 Llama 4 ScoutMeta
19.0%
$0.18 / $0.59
99 Phi-4Microsoft
19.0%
$0.07 / $0.14
99 o4-mini (high reasoning)OpenAI · best of 3 settings
19.0%
$1.10 / $4.40
99 Grok 4 Fast (non-reasoning)xAI
19.0%
–
111 Jamba Large 1.7AI21 Labs
20.0%
–
111 Qwen2.5-1.5B-InstructAlibaba
20.0%
–
111 Qwen2.5-VL-3B-InstructAlibaba
20.0%
–
111 Qwen3-14B (no reasoning)Alibaba · best of 2 settings
20.0%
$0.12 / $0.24
111 Claude 3 OpusAnthropic
20.0%
–
111 Gemini 2.5 Flash Preview (09-2025, no reasoning)Google · best of 2 settings
20.0%
–
111 Gemma 2 27BGoogle
20.0%
$0.65 / $0.65
111 Llama 3.1 8B InstructMeta
20.0%
$0.05 / $0.08
111 Mistral Medium 3Mistral AI
20.0%
$0.40 / $2
111 Kimi K2 ThinkingMoonshot AI
20.0%
$0.60 / $2.50
111 Kimi Linear 48B A3B InstructMoonshot AI
20.0%
–
111 Kimi-VL-A3B-InstructMoonshot AI
20.0%
–
111 Falcon-H1 1.5B Deep InstructTII
20.0%
–
124 Qwen2.5-72B-InstructAlibaba
21.0%
$0.36 / $0.40
124 Command ACohere
21.0%
$2.50 / $10
124 Mixtral 8x22B InstructMistral AI
21.0%
$2 / $6
124 Grok 4.3xAI
21.0%
$1.25 / $2.50
124 Grok 4.5xAI
21.0%
$2 / $6
129 Llama 3.2 1B InstructMeta
22.0%
$0.027 / $0.20
129 Mistral Small 4 (no reasoning)Mistral AI · best of 2 settings
22.0%
$0.15 / $0.60
129 Kimi K2 (0905)Moonshot AI
22.0%
$0.60 / $2.50
129 Kimi K2 (0711)Moonshot AI
22.0%
$0.57 / $2.30
129 GPT-5.6 Terra (extra-high reasoning)OpenAI · best of 5 settings
22.0%
$2 / $12
134 Qwen3.5-397B-A17B (no reasoning)Alibaba
23.0%
$0.55 / $3.50
134 Qwen3-MaxAlibaba
23.0%
$0.78 / $3.90
134 Ministral 3 8BMistral AI · best of 2 settings
23.0%
$0.15 / $0.15
134 Falcon-H1 1.5B InstructTII
23.0%
–
134 Grok 4 Fast (reasoning)xAI
23.0%
–
134 GLM-4-32B-0414Z.ai
23.0%
–
140 Qwen3.5-0.8B (no reasoning)Alibaba
24.0%
–
140 Trinity Large PreviewArcee AI
24.0%
–
140 Gemma 2 9BGoogle
24.0%
–
140 Ministral 3 14BMistral AI · best of 2 settings
24.0%
$0.20 / $0.20
140 GPT-5.2 ChatOpenAI
24.0%
–
140 o1 (high reasoning)OpenAI · best of 2 settings
24.0%
$15 / $60
140 Solar 10.7B Instruct v1.0Upstage
24.0%
–
147 DeepSeek-V3.2-Exp (reasoning on)DeepSeek · best of 2 settings
25.0%
$0.27 / $0.41
147 DeepSeek-V3.2-SpecialeDeepSeek
25.0%
–
147 LFM2-8B-A1BLiquid AI
25.0%
–
147 Grok 4.6xAI
25.0%
$2 / $6
151 Qwen3-1.7B (reasoning on)Alibaba · best of 2 settings
26.0%
–
151 Qwen3-VL-32B-InstructAlibaba
26.0%
$0.10 / $0.42
151 DeepSeek-V3.2 (no reasoning)DeepSeek · best of 2 settings
26.0%
$0.30 / $0.96
151 Devstral Small 2Mistral AI
26.0%
–
151 InternLM3 8B InstructShanghai AI Lab
26.0%
–
156 Mistral Medium 3.1Mistral AI
27.0%
$0.40 / $2
156 Mixtral 8x7B Instruct v0.1Mistral AI
27.0%
–
158 Qwen3-Omni-30B-A3B-ThinkingAlibaba
28.0%
–
158 Qwen3-VL-8B-InstructAlibaba
28.0%
$0.12 / $0.46
158 Command R+ (08-2024)Cohere
28.0%
$2.50 / $10
158 Nemotron Nano 12B v2 VL (BF16, no reasoning)NVIDIA · best of 2 settings
28.0%
–
158 Grok 4.7xAI
28.0%
$2 / $6
158 GLM-4.7 (no reasoning)Z.ai · best of 2 settings
28.0%
$0.54 / $1.98
164 Qwen2.5-VL-72B-InstructAlibaba
29.0%
$0.80 / $1
164 Llama 3.3 8B InstructAllura Forge
29.0%
–
166 Qwen3-235B-A22B-Thinking-2507Alibaba
30.0%
$0.30 / $3
166 Qwen3-VL-235B-A22B-ThinkingAlibaba
30.0%
$0.40 / $4
166 DeepSeek-V3.1 (no reasoning)DeepSeek · best of 2 settings
30.0%
$0.55 / $1.65
166 Ministral 3 8B ReasoningMistral AI · best of 3 settings
30.0%
–
166 GPT-5.4 (extra-high reasoning)OpenAI · best of 5 settings
30.0%
$2.50 / $15
171 Command R (08-2024)Cohere
31.0%
$0.15 / $0.60
171 Llama 3.2 3B InstructMeta
31.0%
$0.05 / $0.33
171 MiniMax-M2.1MiniMax
31.0%
$0.30 / $1.20
174 DeepSeek-V3-0324DeepSeek
32.0%
$0.25 / $1
174 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek · best of 2 settings
32.0%
$0.27 / $1
176 Gemini 3 Pro Preview (low reasoning)Google · best of 2 settings
33.0%
–
177 Qwen2.5-Coder-7B-InstructAlibaba
34.0%
–
177 DeepSeek-V4-Pro (0423, reasoning on)DeepSeek · best of 2 settings
34.0%
$1.42 / $2.83
177 GLM-5 (reasoning on)Z.ai · best of 2 settings
34.0%
$0.95 / $2.55
180 Nemotron 3 Nano 30B A3BNVIDIA · best of 2 settings
35.0%
$0.05 / $0.20
180 GLM-4.6V (reasoning on)Z.ai · best of 2 settings
35.0%
$0.30 / $0.90
182 Jamba Mini 1.7AI21 Labs
36.0%
–
182 MiniMax-M2.7MiniMax
36.0%
$0.30 / $1.20
182 Grok 4.20 (0309, non-reasoning)xAI
36.0%
–
185 Grok 4.20 Beta (0309, non-reasoning)xAI
37.0%
–
186 Inflection 3 ProductivityInflection AI
38.0%
–
186 Grok 4.20 Multi-Agent Beta (0309, 4 agents)xAI
38.0%
–
188 GLM-4.7-Flash (no reasoning)Z.ai · best of 2 settings
39.0%
$0.06 / $0.40
189 Qwen3-VL-4B-InstructAlibaba
40.0%
–
189 Llama 2 70B ChatMeta
40.0%
–
189 Falcon-H1 0.5B InstructTII
40.0%
–
192 MiniMax-M2.5MiniMax
41.0%
$0.30 / $1.20
193 Ring-1TAnt Group
42.0%
–
193 Inflection 3 PiInflection AI
42.0%
–
193 Step 3.5 FlashStepFun
42.0%
$0.10 / $0.30
193 MiMo-V2-Flash (reasoning on)Xiaomi · best of 2 settings
42.0%
–
197 Rnj-1 InstructEssential AI
43.0%
–
197 Grok 4.20 Beta (0309, reasoning)xAI
43.0%
–
199 Ministral 3 14B ReasoningMistral AI · best of 2 settings
44.0%
–
200 Nova 2 Lite (reasoning on)Amazon · best of 2 settings
45.0%
$0.30 / $2.50
200 DeepSeek-R1-0528DeepSeek
45.0%
$0.50 / $2.18
202 Grok 4.20 (0309, reasoning)xAI
46.0%
–
203 Granite 4.0 H SmallIBM
47.0%
–
204 EXAONE 4.0 32BLG AI Research
50.0%
–
205 Reka Flash 3Reka AI
54.0%
$0.10 / $0.20
206 Apriel Nemotron 15B ThinkerServiceNow
56.0%
–
207 GPT-5 (high reasoning)OpenAI · best of 2 settings
58.0%
$1.25 / $10
208 Seed-OSS-36B-Instruct (no reasoning)ByteDance · best of 3 settings
59.0%
–
209 Qwen2.5-VL-32B-InstructAlibaba
64.0%
–
209 Qwen3-4B-Thinking-2507Alibaba
64.0%
–
211 Qwen3-Next-80B-A3B-ThinkingAlibaba
67.0%
$0.15 / $1.20
212 Kimi-VL-A3B-Thinking (2506)Moonshot AI
71.0%
–
213 MiniMax-M2MiniMax
73.0%
$0.30 / $1.20
214 GPT-5.2 (high reasoning)OpenAI · best of 4 settings
76.0%
$1.75 / $14
215 Qwen3-VL-4B-ThinkingAlibaba
79.0%
–
216 Qwen3-30B-A3B-Thinking-2507Alibaba
82.0%
$0.20 / $2.40
216 QwQ-32BAlibaba
82.0%
–
218 Qwen3-VL-2B-ThinkingAlibaba
87.0%
–
219 Qwen3-VL-8B-ThinkingAlibaba
89.0%
$0.18 / $2.10
220 Qwen3-VL-32B-ThinkingAlibaba
91.0%
–
221 gpt-oss-20b (low reasoning)OpenAI · best of 3 settings
98.0%
$0.03 / $0.15
222 Olmo 3 32B ThinkAi2
102.0%
–
223 Nanbeige4-3B-Thinking (2510)Nanbeige
103.0%
–
224 gpt-oss-120b (medium reasoning)OpenAI
107.0%
$0.15 / $0.60
225 Solar Pro 3Upstage
108.0%
$0.15 / $0.60
226 Qwen3-VL-30B-A3B-ThinkingAlibaba
113.0%
$0.29 / $1
226 GPT-5.1 (high reasoning)OpenAI · best of 3 settings
113.0%
$1.25 / $10
228 Olmo 3 7B ThinkAi2
117.0%
–
229 Ling-1TAnt Group
120.0%
–
230 Nanbeige4-3B-Thinking (2511)Nanbeige
173.0%
–

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Results as published by UGI Leaderboard; we do not re-run them.

What it measures

Average percentage difference between a requested and the delivered word count.

What it does not measure

Not other format limits such as character counts or bullet counts.

947 results from UGI Leaderboard not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Community fine-tunes and merges: 934
  • No score published: 13

UGI Leaderboard by DontPlanToEnd, Apache License 2.0.