Benchmarks / UGI Leaderboard

Reported by UGI Leaderboard

UGI Leaderboard

Average percentage difference between a requested and the delivered word count.

Last updated 2 Oct 2026

Results dated
6 Sep 2025 to 2 Oct 2026
Results
379 configurations of 230 models
Unit
% off the requested word count
Licence
Apache License 2.0

Requested-length error: Claude Opus 4.7

375 results above zero · % off the requested word count, lower is better. Choose a model to highlight it.Clear highlight

  1. 5 Claude Opus 4.7 (low reasoning)Anthropic 1.0%
  2. 5 GPT-5.6 Sol (extra-high reasoning)OpenAI 1.0%
  3. 5 GPT-6.1 Sol (low reasoning)OpenAI 1.0%
  4. 5 GPT-6 Astra (low reasoning)OpenAI 1.0%
  5. 5 GPT-6 Astra (medium reasoning)OpenAI 1.0%
  6. 10 Claude Fable 5 (high reasoning)Anthropic 2.0%
  7. 10 Claude Fable 5 (low reasoning)Anthropic 2.0%
  8. 10 Claude Fable 5 (max reasoning)Anthropic 2.0%
  9. 10 Claude Fable 5 (medium reasoning)Anthropic 2.0%
  10. 10 Claude Fable 5 (extra-high reasoning)Anthropic 2.0%
  11. 10 Claude Opus 4.7 (high reasoning)Anthropic 2.0%
  12. 10 Claude Opus 4.7 (medium reasoning)Anthropic 2.0%
  13. 10 Claude Opus 5.5 (max reasoning)Anthropic 2.0%
  14. 10 Claude Sonnet 4.6 (low reasoning)Anthropic 2.0%
  15. 10 Claude Sonnet 4.6 (max reasoning)Anthropic 2.0%
  16. 22 Claude Opus 4.7 (max reasoning)Anthropic 3.0%

4 models had none: GPT-6 Astra (extra high), GPT-6 Astra (high reasoning), GPT-6.1 Sol (extra high), GPT-6.1 Sol (high reasoning).

Full results

UGI: requested-length error, % off the requested word count, lower is better
#ModelRequested-length error
% off the requested word count, lower is better
Price
$ per million tokens, in / out
1≈ GPT-6.1 Sol (high reasoning)OpenAI
0.0%
$2 / $10
1≈ GPT-6.1 Sol (extra-high reasoning)OpenAI
0.0%
$2 / $10
1≈ GPT-6 Astra (high reasoning)OpenAI
0.0%
$10 / $50
1≈ GPT-6 Astra (extra-high reasoning)OpenAI
0.0%
$10 / $50
5 Claude Opus 4.7 (low reasoning)Anthropic
1.0%
$5 / $25
5 GPT-5.6 Sol (extra-high reasoning)OpenAI
1.0%
$4 / $20
5 GPT-6.1 Sol (low reasoning)OpenAI
1.0%
$2 / $10
5 GPT-6 Astra (low reasoning)OpenAI
1.0%
$10 / $50
5 GPT-6 Astra (medium reasoning)OpenAI
1.0%
$10 / $50
10 Claude Fable 5 (high reasoning)Anthropic
2.0%
$10 / $50
10 Claude Fable 5 (low reasoning)Anthropic
2.0%
$10 / $50
10 Claude Fable 5 (max reasoning)Anthropic
2.0%
$10 / $50
10 Claude Fable 5 (medium reasoning)Anthropic
2.0%
$10 / $50
10 Claude Fable 5 (extra-high reasoning)Anthropic
2.0%
$10 / $50
10 Claude Opus 4.7 (high reasoning)Anthropic
2.0%
$5 / $25
10 Claude Opus 4.7 (medium reasoning)Anthropic
2.0%
$5 / $25
10 Claude Opus 5.5 (max reasoning)Anthropic
2.0%
$4 / $20
10 Claude Sonnet 4.6 (low reasoning)Anthropic
2.0%
$3 / $15
10 Claude Sonnet 4.6 (max reasoning)Anthropic
2.0%
$3 / $15
10 Claude Sonnet 4.6 (medium reasoning)Anthropic
2.0%
$3 / $15
10 GPT-6.1 Sol (medium reasoning)OpenAI
2.0%
$2 / $10
22 Claude Fable 5.1 (high reasoning)Anthropic
3.0%
$10 / $50
22 Claude Fable 5.1 (medium reasoning)Anthropic
3.0%
$10 / $50
22 Claude Opus 4.1 (no reasoning)Anthropic
3.0%
$15 / $75
22 Claude Opus 4.5 (no reasoning)Anthropic
3.0%
$5 / $25
22 Claude Opus 4.6 (max reasoning)Anthropic
3.0%
$5 / $25
22 Claude Opus 4.7 (max reasoning)Anthropic
3.0%
$5 / $25
22 Claude Sonnet 4.6 (high reasoning)Anthropic
3.0%
$3 / $15
22 Claude Sonnet 5 (high reasoning)Anthropic
3.0%
$2 / $10
22 Claude Sonnet 5 (low reasoning)Anthropic
3.0%
$2 / $10
22 Claude Sonnet 5 (medium reasoning)Anthropic
3.0%
$2 / $10
32 Claude Fable 5.1 (low reasoning)Anthropic
4.0%
$10 / $50
32 Claude Opus 4.5 (reasoning on)Anthropic
4.0%
$5 / $25
32 Claude Opus 4.6 (low reasoning)Anthropic
4.0%
$5 / $25
32 Claude Opus 4.6 (medium reasoning)Anthropic
4.0%
$5 / $25
32 Claude Sonnet 4 (no reasoning)Anthropic
4.0%
$3 / $15
32 GPT-5.6 Sol (high reasoning)OpenAI
4.0%
$4 / $20
32 GPT-6 Sol (high reasoning)OpenAI
4.0%
$2 / $10
32 GPT-6 Sol (medium reasoning)OpenAI
4.0%
$2 / $10
32 GPT-6 Sol (extra-high reasoning)OpenAI
4.0%
$2 / $10
41 Claude Opus 4.6 (high reasoning)Anthropic
5.0%
$5 / $25
41 Claude Opus 5.5 (high reasoning)Anthropic
5.0%
$4 / $20
41 Claude Opus 5.5 (low reasoning)Anthropic
5.0%
$4 / $20
41 Claude Opus 5.5 (medium reasoning)Anthropic
5.0%
$4 / $20
41 Claude Opus 5.5 (extra-high reasoning)Anthropic
5.0%
$4 / $20
41 Claude Sonnet 4.5 (no reasoning)Anthropic
5.0%
$3 / $15
41 Gemma 4 31BGoogle
5.0%
$0.14 / $0.40
41 GPT-6 Sol (low reasoning)OpenAI
5.0%
$2 / $10
49 Claude Haiku 4.5 (no reasoning)Anthropic
6.0%
$1 / $5
49 Claude Opus 4.8 (extra-high reasoning)Anthropic
6.0%
$5 / $25
49 Claude Opus 4 (no reasoning)Anthropic
6.0%
–
49 Claude Opus 4 (reasoning on)Anthropic
6.0%
–
49 Claude Sonnet 5 (max reasoning)Anthropic
6.0%
$2 / $10
49 Gemma 4 31B (reasoning by prefill)Google
6.0%
$0.14 / $0.40
49 GPT-4.1OpenAI
6.0%
$2 / $8
56 Qwen3-Next-80B-A3B-InstructAlibaba
7.0%
$0.10 / $1.10
56 Claude 3.7 Sonnet (no reasoning)Anthropic
7.0%
–
56 Claude Opus 4.8 (low reasoning)Anthropic
7.0%
$5 / $25
56 Claude Sonnet 5 (extra-high reasoning)Anthropic
7.0%
$2 / $10
56 Gemini 3.7 Flash (medium reasoning)Google
7.0%
$1.50 / $7.50
56 GPT-5.6 Sol (medium reasoning)OpenAI
7.0%
$4 / $20
56 Grok 4.1 Fast (reasoning)xAI
7.0%
–
63 Qwen3-4B-Instruct-2507Alibaba
8.0%
–
63 Qwen3.6-Plus (reasoning on)Alibaba
8.0%
$0.33 / $1.95
63 Claude Sonnet 4.5 (reasoning on)Anthropic
8.0%
$3 / $15
63 Gemini 3.7 Flash (high reasoning)Google
8.0%
$1.50 / $7.50
63 Gemma 3 27BGoogle
8.0%
$0.12 / $0.20
63 Gemma 4 12B (reasoning by prefill)Google
8.0%
–
63 Gemma 4 12BGoogle
8.0%
–
63 Gemma 4 26B A4B (reasoning by prefill)Google
8.0%
$0.10 / $0.30
63 Gemma 4 26B A4BGoogle
8.0%
$0.10 / $0.30
63 Mistral Small 3Mistral AI
8.0%
$0.05 / $0.08
63 GPT-5.3 ChatOpenAI
8.0%
–
63 MiMo-V2.5-Pro (reasoning on)Xiaomi
8.0%
$0.43 / $0.87
75 Qwen3-4B (no reasoning)Alibaba
9.0%
–
75 Qwen3.5-122B-A10B (reasoning on)Alibaba
9.0%
$0.26 / $2.08
75 Qwen3.6-27B (no reasoning)Alibaba
9.0%
$0.30 / $3.20
75 Claude Opus 4.1 (reasoning on)Anthropic
9.0%
$15 / $75
75 Claude Opus 4.8 (high reasoning)Anthropic
9.0%
$5 / $25
75 Claude Opus 4.8 (max reasoning)Anthropic
9.0%
$5 / $25
75 Gemini 3.5 Flash (low reasoning)Google
9.0%
$1.50 / $9
75 Gemini 3.6 Flash (minimal reasoning)Google
9.0%
$1.50 / $7.50
75 Gemini 3.7 Flash (low reasoning)Google
9.0%
$1.50 / $7.50
75 Gemma 3 12BGoogle
9.0%
$0.05 / $0.15
75 Gemma 3 4BGoogle
9.0%
$0.05 / $0.10
75 Gemma 4 E2BGoogle
9.0%
–
75 ChatGPT-4o (2025-03-26)OpenAI
9.0%
–
75 GLM-5.1 (reasoning on)Z.ai
9.0%
$1.38 / $4.40
89 Qwen3.6-35B-A3B (no reasoning)Alibaba
10.0%
$0.10 / $1
89 Claude 3.7 Sonnet (reasoning on)Anthropic
10.0%
–
89 Claude Opus 4.8 (medium reasoning)Anthropic
10.0%
$5 / $25
89 Gemma 4 E2B (reasoning by prefill)Google
10.0%
–
89 Gemma 4 E4B (reasoning by prefill)Google
10.0%
–
89 Llama 3.3 70B InstructMeta
10.0%
$0.59 / $0.79
89 Kimi K2.5 (reasoning on)Moonshot AI
10.0%
$0.57 / $2.85
89 GPT-5.6 Sol (low reasoning)OpenAI
10.0%
$4 / $20
89 Grok 4.1 Fast (non-reasoning)xAI
10.0%
–
89 GLM-4.5-Air (no reasoning)Z.ai
10.0%
$0.14 / $0.86
99 Qwen3.5-27B (reasoning by prefill)Alibaba
11.0%
$0.27 / $2.16
99 Qwen3.6-Plus (no reasoning)Alibaba
11.0%
$0.33 / $1.95
99 Gemini 3.6 Flash (high reasoning)Google
11.0%
$1.50 / $7.50
99 Mistral Large 2 (2407)Mistral AI
11.0%
$2 / $6
99 Mistral Large 2.1 (2411)Mistral AI
11.0%
–
99 Mistral Small 3.1 24BMistral AI
11.0%
$0.35 / $0.56
99 Mistral Small (2409)Mistral AI
11.0%
–
106 DeepSeek-V4-Flash (0423, reasoning on)DeepSeek
12.0%
$0.14 / $0.28
106 Gemini 3.5 Flash (medium reasoning)Google
12.0%
$1.50 / $9
106 Gemini 3.6 Flash (medium reasoning)Google
12.0%
$1.50 / $7.50
106 GPT-4o (2024-05-13)OpenAI
12.0%
$5 / $15
106 GLM-5.2 (no reasoning)Z.ai
12.0%
$1.40 / $4.40
111 Qwen3-30B-A3B-Instruct-2507Alibaba
13.0%
$0.09 / $0.30
111 Claude Sonnet 4 (reasoning on)Anthropic
13.0%
$3 / $15
111 Gemini 3.5 Flash (minimal reasoning)Google
13.0%
$1.50 / $9
111 Gemini 3.8 Flash (high reasoning)Google
13.0%
$1.50 / $7.50
111 Llama 4 MaverickMeta
13.0%
$0.27 / $0.85
111 GPT-5.5 (extra-high reasoning)OpenAI
13.0%
$5 / $30
111 GLM-4.5 (no reasoning)Z.ai
13.0%
$0.60 / $2.20
111 GLM-4.6 (reasoning on)Z.ai
13.0%
$0.50 / $2
119 Qwen3-30B-A3B (no reasoning)Alibaba
14.0%
$0.12 / $0.50
119 Gemini 3.5 Flash (high reasoning)Google
14.0%
$1.50 / $9
119 Gemini 3.8 Flash (medium reasoning)Google
14.0%
$1.50 / $7.50
119 MedGemma 27B TextGoogle
14.0%
–
119 LFM2-24B-A2BLiquid AI
14.0%
–
119 Llama 3.1 405B InstructMeta
14.0%
–
119 Muse Glimmer 30B (extra-high reasoning)Meta
14.0%
$0.30 / $1.20
119 GPT-5.1 ChatOpenAI
14.0%
–
127 Qwen3-235B-A22B-Instruct-2507Alibaba
15.0%
$0.15 / $0.75
127 Qwen3.5-9B (reasoning by prefill)Alibaba
15.0%
$0.10 / $0.15
127 Olmo 3 7B InstructAi2
15.0%
–
127 Claude 3 HaikuAnthropic
15.0%
–
127 Gemma 4 E4BGoogle
15.0%
–
127 Llama 3.1 70B InstructMeta
15.0%
$0.40 / $0.40
127 Mistral Large 3 (no reasoning)Mistral AI
15.0%
$0.50 / $1.50
127 GPT-5.6 Luna (extra-high reasoning)OpenAI
15.0%
$0.20 / $1.20
135 Qwen3.5-2B (no reasoning)Alibaba
16.0%
–
135 Qwen3.5-4B (reasoning by prefill)Alibaba
16.0%
–
135 Qwen3.6-35B-A3B (reasoning by prefill)Alibaba
16.0%
$0.10 / $1
135 Qwen3-Coder-30B-A3B-InstructAlibaba
16.0%
$0.07 / $0.28
135 Gemini 3.8 Flash (low reasoning)Google
16.0%
$1.50 / $7.50
135 Gemini 3 Flash Preview (minimal reasoning)Google
16.0%
$0.50 / $3
135 Muse Glimmer 30B (high reasoning)Meta
16.0%
$0.30 / $1.20
135 GPT-5.6 Luna (high reasoning)OpenAI
16.0%
$0.20 / $1.20
135 GPT-5.6 Sol (no reasoning)OpenAI
16.0%
$4 / $20
135 GPT-5 Chat (latest)OpenAI
16.0%
–
135 Grok 3xAI
16.0%
–
135 MiMo-V2.5 (reasoning on)Xiaomi
16.0%
$0.17 / $0.34
135 GLM-5.1 (no reasoning)Z.ai
16.0%
$1.38 / $4.40
148 Qwen2.5-32B-InstructAlibaba
17.0%
–
148 Qwen3.5-2B (reasoning by prefill)Alibaba
17.0%
–
148 Qwen3.5-35B-A3B (no reasoning)Alibaba
17.0%
$0.16 / $1.30
148 Qwen3-8B (no reasoning)Alibaba
17.0%
$0.12 / $0.46
148 DeepSeek-V3DeepSeek
17.0%
$0.26 / $1.03
148 Magistral Small 1.2Mistral AI
17.0%
–
148 Mistral Small 3.2 24BMistral AI
17.0%
$0.094 / $0.25
148 GPT-5.6 Luna (no reasoning)OpenAI
17.0%
$0.20 / $1.20
148 Falcon-H1 7B InstructTII
17.0%
–
148 Grok 4 (0709)xAI
17.0%
–
158 Qwen3-0.6B (reasoning on)Alibaba
18.0%
–
158 Qwen3.5-35B-A3B (reasoning by prefill)Alibaba
18.0%
$0.16 / $1.30
158 Qwen3.5-4B (no reasoning)Alibaba
18.0%
–
158 DeepSeek-R1DeepSeek
18.0%
$0.70 / $2.50
158 Mistral Medium 3.5 (no reasoning)Mistral AI
18.0%
$1.50 / $7.50
158 Mistral NemoMistral AI
18.0%
$0.023 / $0.03
158 Kimi K2.6 (reasoning on)Moonshot AI
18.0%
$0.95 / $4
158 o3 (high reasoning)OpenAI
18.0%
$2 / $8
158 o3 (medium reasoning)OpenAI
18.0%
$2 / $8
158 Hy3 Preview (reasoning on)Tencent
18.0%
$0.18 / $0.60
158 Falcon-H1 3B InstructTII
18.0%
–
158 GLM-4.6 (no reasoning)Z.ai
18.0%
$0.50 / $2
158 GLM-5.2 (reasoning on)Z.ai
18.0%
$1.40 / $4.40
171 Qwen2.5-7B-InstructAlibaba
19.0%
$0.10 / $0.20
171 Qwen2.5-14B-InstructAlibaba
19.0%
–
171 Qwen3-32B (no reasoning)Alibaba
19.0%
$0.14 / $0.40
171 Qwen3.5-27B (no reasoning)Alibaba
19.0%
$0.27 / $2.16
171 Qwen3-VL-235B-A22B-InstructAlibaba
19.0%
$0.30 / $1.50
171 Qwen3-VL-2B-InstructAlibaba
19.0%
–
171 DeepSeek-V4-Flash (0423, no reasoning)DeepSeek
19.0%
$0.14 / $0.28
171 Gemini 2.5 ProGoogle
19.0%
$1.25 / $10
171 Gemini 3.1 Pro Preview (low reasoning)Google
19.0%
$2 / $12
171 Gemini 3.1 Pro Preview (medium reasoning)Google
19.0%
$2 / $12
171 Gemma 2 2BGoogle
19.0%
–
171 Llama 4 ScoutMeta
19.0%
$0.18 / $0.59
171 Phi-4Microsoft
19.0%
$0.07 / $0.14
171 GPT-5.6 Luna (low reasoning)OpenAI
19.0%
$0.20 / $1.20
171 o4-mini (high reasoning)OpenAI
19.0%
$1.10 / $4.40
171 Grok 4 Fast (non-reasoning)xAI
19.0%
–
187 Jamba Large 1.7AI21 Labs
20.0%
–
187 Qwen2.5-1.5B-InstructAlibaba
20.0%
–
187 Qwen2.5-VL-3B-InstructAlibaba
20.0%
–
187 Qwen3-14B (no reasoning)Alibaba
20.0%
$0.12 / $0.24
187 Qwen3.5-122B-A10B (no reasoning)Alibaba
20.0%
$0.26 / $2.08
187 Claude 3 OpusAnthropic
20.0%
–
187 Gemini 2.5 Flash Preview (09-2025, no reasoning)Google
20.0%
–
187 Gemini 3.1 Pro Preview (high reasoning)Google
20.0%
$2 / $12
187 Gemma 2 27BGoogle
20.0%
$0.65 / $0.65
187 Llama 3.1 8B InstructMeta
20.0%
$0.05 / $0.08
187 Mistral Medium 3Mistral AI
20.0%
$0.40 / $2
187 Kimi K2 ThinkingMoonshot AI
20.0%
$0.60 / $2.50
187 Kimi Linear 48B A3B InstructMoonshot AI
20.0%
–
187 Kimi-VL-A3B-InstructMoonshot AI
20.0%
–
187 GPT-5.6 Luna (medium reasoning)OpenAI
20.0%
$0.20 / $1.20
187 Falcon-H1 1.5B Deep InstructTII
20.0%
–
203 Qwen2.5-72B-InstructAlibaba
21.0%
$0.36 / $0.40
203 Qwen3.5-9B (no reasoning)Alibaba
21.0%
$0.10 / $0.15
203 Command ACohere
21.0%
$2.50 / $10
203 Mixtral 8x22B InstructMistral AI
21.0%
$2 / $6
203 Grok 4.3xAI
21.0%
$1.25 / $2.50
203 Grok 4.5xAI
21.0%
$2 / $6
203 GLM-4.5 (reasoning on)Z.ai
21.0%
$0.60 / $2.20
210 Llama 3.2 1B InstructMeta
22.0%
$0.027 / $0.20
210 Muse Glimmer 30B (medium reasoning)Meta
22.0%
$0.30 / $1.20
210 Mistral Small 4 (no reasoning)Mistral AI
22.0%
$0.15 / $0.60
210 Kimi K2 (0905)Moonshot AI
22.0%
$0.60 / $2.50
210 Kimi K2 (0711)Moonshot AI
22.0%
$0.57 / $2.30
210 GPT-5.6 Terra (extra-high reasoning)OpenAI
22.0%
$2 / $12
216 Qwen3.5-397B-A17B (no reasoning)Alibaba
23.0%
$0.55 / $3.50
216 Qwen3-MaxAlibaba
23.0%
$0.78 / $3.90
216 Gemini 3 Flash Preview (medium reasoning)Google
23.0%
$0.50 / $3
216 Ministral 3 8BMistral AI
23.0%
$0.15 / $0.15
216 Mistral Large 3 (reasoning on)Mistral AI
23.0%
$0.50 / $1.50
216 Falcon-H1 1.5B InstructTII
23.0%
–
216 Grok 4 Fast (reasoning)xAI
23.0%
–
216 MiMo-V2.5-Pro (no reasoning)Xiaomi
23.0%
$0.43 / $0.87
216 GLM-4-32B-0414Z.ai
23.0%
–
225 Qwen3.5-0.8B (no reasoning)Alibaba
24.0%
–
225 Qwen3.6-27B (reasoning by prefill)Alibaba
24.0%
$0.30 / $3.20
225 Claude Haiku 4.5 (reasoning on)Anthropic
24.0%
$1 / $5
225 Trinity Large PreviewArcee AI
24.0%
–
225 Gemini 2.5 Flash Preview (09-2025, reasoning on)Google
24.0%
–
225 Gemma 2 9BGoogle
24.0%
–
225 Ministral 3 14BMistral AI
24.0%
$0.20 / $0.20
225 Kimi K2.5 (no reasoning)Moonshot AI
24.0%
$0.57 / $2.85
225 GPT-5.2 ChatOpenAI
24.0%
–
225 o1 (high reasoning)OpenAI
24.0%
$15 / $60
225 Solar 10.7B Instruct v1.0Upstage
24.0%
–
236 DeepSeek-V3.2-Exp (reasoning on)DeepSeek
25.0%
$0.27 / $0.41
236 DeepSeek-V3.2-SpecialeDeepSeek
25.0%
–
236 LFM2-8B-A1BLiquid AI
25.0%
–
236 Muse Glimmer 30B (low reasoning)Meta
25.0%
$0.30 / $1.20
236 Grok 4.6xAI
25.0%
$2 / $6
241 Qwen3-1.7B (reasoning on)Alibaba
26.0%
–
241 Qwen3-VL-32B-InstructAlibaba
26.0%
$0.10 / $0.42
241 DeepSeek-V3.2 (no reasoning)DeepSeek
26.0%
$0.30 / $0.96
241 Gemini 3 Flash Preview (high reasoning)Google
26.0%
$0.50 / $3
241 Devstral Small 2Mistral AI
26.0%
–
241 InternLM3 8B InstructShanghai AI Lab
26.0%
–
247 DeepSeek-V3.2-Exp (no reasoning)DeepSeek
27.0%
$0.27 / $0.41
247 Mistral Medium 3.1Mistral AI
27.0%
$0.40 / $2
247 Mixtral 8x7B Instruct v0.1Mistral AI
27.0%
–
247 o4-mini (medium reasoning)OpenAI
27.0%
$1.10 / $4.40
251 Qwen3-Omni-30B-A3B-ThinkingAlibaba
28.0%
–
251 Qwen3-VL-8B-InstructAlibaba
28.0%
$0.12 / $0.46
251 Command R+ (08-2024)Cohere
28.0%
$2.50 / $10
251 Nemotron Nano 12B v2 VL (BF16, no reasoning)NVIDIA
28.0%
–
251 Grok 4.7xAI
28.0%
$2 / $6
251 GLM-4.7 (no reasoning)Z.ai
28.0%
$0.54 / $1.98
257 Qwen2.5-VL-72B-InstructAlibaba
29.0%
$0.80 / $1
257 Llama 3.3 8B InstructAllura Forge
29.0%
–
257 DeepSeek-V3.2 (reasoning on)DeepSeek
29.0%
$0.30 / $0.96
257 o3 (low reasoning)OpenAI
29.0%
$2 / $8
257 MiMo-V2.5 (no reasoning)Xiaomi
29.0%
$0.17 / $0.34
262 Qwen3-235B-A22B-Thinking-2507Alibaba
30.0%
$0.30 / $3
262 Qwen3-VL-235B-A22B-ThinkingAlibaba
30.0%
$0.40 / $4
262 DeepSeek-V3.1 (no reasoning)DeepSeek
30.0%
$0.55 / $1.65
262 Ministral 3 8B ReasoningMistral AI
30.0%
–
262 GPT-5.4 (extra-high reasoning)OpenAI
30.0%
$2.50 / $15
267 Command R (08-2024)Cohere
31.0%
$0.15 / $0.60
267 Llama 3.2 3B InstructMeta
31.0%
$0.05 / $0.33
267 MiniMax-M2.1MiniMax
31.0%
$0.30 / $1.20
267 GPT-5.6 Terra (high reasoning)OpenAI
31.0%
$2 / $12
267 GPT-5.6 Terra (medium reasoning)OpenAI
31.0%
$2 / $12
272 Qwen3-1.7B (no reasoning)Alibaba
32.0%
–
272 DeepSeek-V3-0324DeepSeek
32.0%
$0.25 / $1
272 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek
32.0%
$0.27 / $1
272 Hy3 Preview (no reasoning)Tencent
32.0%
$0.18 / $0.60
276 Gemini 3 Pro Preview (low reasoning)Google
33.0%
–
276 GLM-4.7 (reasoning on)Z.ai
33.0%
$0.54 / $1.98
278 Qwen2.5-Coder-7B-InstructAlibaba
34.0%
–
278 DeepSeek-V3.1 (reasoning on)DeepSeek
34.0%
$0.55 / $1.65
278 DeepSeek-V3.1-Terminus (no reasoning)DeepSeek
34.0%
$0.27 / $1
278 DeepSeek-V4-Pro (0423, reasoning on)DeepSeek
34.0%
$1.42 / $2.83
278 Gemini 3 Pro Preview (high reasoning)Google
34.0%
–
278 Mistral Small 4 (high reasoning)Mistral AI
34.0%
$0.15 / $0.60
278 Kimi K2.6 (no reasoning)Moonshot AI
34.0%
$0.95 / $4
278 GLM-5 (reasoning on)Z.ai
34.0%
$0.95 / $2.55
286 Ministral 3 8B (reasoning by prefill)Mistral AI
35.0%
$0.15 / $0.15
286 Nemotron 3 Nano 30B A3BNVIDIA
35.0%
$0.05 / $0.20
286 GPT-5.6 Terra (low reasoning)OpenAI
35.0%
$2 / $12
286 GLM-4.6V (reasoning on)Z.ai
35.0%
$0.30 / $0.90
290 Jamba Mini 1.7AI21 Labs
36.0%
–
290 MiniMax-M2.7MiniMax
36.0%
$0.30 / $1.20
290 o1 (low reasoning)OpenAI
36.0%
$15 / $60
290 o4-mini (low reasoning)OpenAI
36.0%
$1.10 / $4.40
290 Grok 4.20 (0309, non-reasoning)xAI
36.0%
–
290 GLM-5 (no reasoning)Z.ai
36.0%
$0.95 / $2.55
296 Grok 4.20 Beta (0309, non-reasoning)xAI
37.0%
–
297 Inflection 3 ProductivityInflection AI
38.0%
–
297 Grok 4.20 Multi-Agent Beta (0309, 4 agents)xAI
38.0%
–
299 GLM-4.7-Flash (no reasoning)Z.ai
39.0%
$0.06 / $0.40
300 Qwen3-VL-4B-InstructAlibaba
40.0%
–
300 Llama 2 70B ChatMeta
40.0%
–
300 Falcon-H1 0.5B InstructTII
40.0%
–
303 MiniMax-M2.5MiniMax
41.0%
$0.30 / $1.20
304 Ring-1TAnt Group
42.0%
–
304 DeepSeek-V4-Pro (0423, no reasoning)DeepSeek
42.0%
$1.42 / $2.83
304 Inflection 3 PiInflection AI
42.0%
–
304 GPT-5.5 (high reasoning)OpenAI
42.0%
$5 / $30
304 GPT-5.6 Terra (no reasoning)OpenAI
42.0%
$2 / $12
304 Step 3.5 FlashStepFun
42.0%
$0.10 / $0.30
304 MiMo-V2-Flash (reasoning on)Xiaomi
42.0%
–
311 Rnj-1 InstructEssential AI
43.0%
–
311 Grok 4.20 Beta (0309, reasoning)xAI
43.0%
–
313 Ministral 3 14B ReasoningMistral AI
44.0%
–
313 MiMo-V2-Flash (no reasoning)Xiaomi
44.0%
–
315 Nova 2 Lite (reasoning on)Amazon
45.0%
$0.30 / $2.50
315 DeepSeek-R1-0528DeepSeek
45.0%
$0.50 / $2.18
317 Nova 2 Lite (no reasoning)Amazon
46.0%
$0.30 / $2.50
317 GPT-5.4 (high reasoning)OpenAI
46.0%
$2.50 / $15
317 Grok 4.20 (0309, reasoning)xAI
46.0%
–
320 Granite 4.0 H SmallIBM
47.0%
–
321 EXAONE 4.0 32BLG AI Research
50.0%
–
321 GPT-5.4 (no reasoning)OpenAI
50.0%
$2.50 / $15
323 Qwen3-14B (reasoning on)Alibaba
53.0%
$0.12 / $0.24
323 GPT-5.4 (medium reasoning)OpenAI
53.0%
$2.50 / $15
325 Reka Flash 3Reka AI
54.0%
$0.10 / $0.20
326 Apriel Nemotron 15B ThinkerServiceNow
56.0%
–
327 GPT-5.4 (low reasoning)OpenAI
57.0%
$2.50 / $15
328 Qwen3-30B-A3B (reasoning on)Alibaba
58.0%
$0.12 / $0.50
328 Ministral 3 14B (reasoning by prefill)Mistral AI
58.0%
$0.20 / $0.20
328 GPT-5 (high reasoning)OpenAI
58.0%
$1.25 / $10
331 Seed-OSS-36B-Instruct (no reasoning)ByteDance
59.0%
–
332 Mistral Medium 3.5 (high reasoning)Mistral AI
60.0%
$1.50 / $7.50
333 GLM-4.5-AirZ.ai
61.0%
$0.14 / $0.86
334 Qwen3-4B (reasoning on)Alibaba
62.0%
–
335 Qwen2.5-VL-32B-InstructAlibaba
64.0%
–
335 Qwen3-4B-Thinking-2507Alibaba
64.0%
–
335 Qwen3-8B (reasoning on)Alibaba
64.0%
$0.12 / $0.46
338 Qwen3-Next-80B-A3B-ThinkingAlibaba
67.0%
$0.15 / $1.20
339 Kimi-VL-A3B-Thinking (2506)Moonshot AI
71.0%
–
339 GPT-5.5 (medium reasoning)OpenAI
71.0%
$5 / $30
341 MiniMax-M2MiniMax
73.0%
$0.30 / $1.20
342 Qwen3-32B (reasoning on)Alibaba
74.0%
$0.14 / $0.40
343 GPT-5.2 (high reasoning)OpenAI
76.0%
$1.75 / $14
344 Qwen3-VL-4B-ThinkingAlibaba
79.0%
–
345 Nemotron Nano 12B v2 VL (BF16)NVIDIA
81.0%
–
346 Qwen3-30B-A3B-Thinking-2507Alibaba
82.0%
$0.20 / $2.40
346 QwQ-32BAlibaba
82.0%
–
346 Ministral 3 8B Reasoning (reasoning by prefill)Mistral AI
82.0%
–
349 GLM-4.6V (no reasoning)Z.ai
86.0%
$0.30 / $0.90
350 Qwen3-VL-2B-ThinkingAlibaba
87.0%
–
351 Qwen3-VL-8B-ThinkingAlibaba
89.0%
$0.18 / $2.10
352 Qwen3-VL-32B-ThinkingAlibaba
91.0%
–
353 GPT-5 (low reasoning)OpenAI
92.0%
$1.25 / $10
354 GPT-5.2 (medium reasoning)OpenAI
93.0%
$1.75 / $14
355 GPT-5.5 (low reasoning)OpenAI
96.0%
$5 / $30
356 GPT-5.5 (no reasoning)OpenAI
98.0%
$5 / $30
356 gpt-oss-20b (low reasoning)OpenAI
98.0%
$0.03 / $0.15
358 Olmo 3 32B ThinkAi2
102.0%
–
359 Nanbeige4-3B-Thinking (2510)Nanbeige
103.0%
–
359 GPT-5.2 (low reasoning)OpenAI
103.0%
$1.75 / $14
361 Seed-OSS-36B-Instruct (512-token reasoning budget)ByteDance
104.0%
–
362 gpt-oss-120b (medium reasoning)OpenAI
107.0%
$0.15 / $0.60
363 Solar Pro 3Upstage
108.0%
$0.15 / $0.60
364 GPT-5.2 (no reasoning)OpenAI
112.0%
$1.75 / $14
365 Qwen3-VL-30B-A3B-ThinkingAlibaba
113.0%
$0.29 / $1
365 GPT-5.1 (high reasoning)OpenAI
113.0%
$1.25 / $10
367 Olmo 3 7B ThinkAi2
117.0%
–
368 Ling-1TAnt Group
120.0%
–
369 Ministral 3 14B Reasoning (reasoning by prefill)Mistral AI
123.0%
–
370 GPT-5.1 (medium reasoning)OpenAI
129.0%
$1.25 / $10
371 GPT-5.1 (low reasoning)OpenAI
157.0%
$1.25 / $10
372 Seed-OSS-36B-InstructByteDance
162.0%
–
373 GLM-4.7-Flash (reasoning by prefill)Z.ai
172.0%
$0.06 / $0.40
374 Nanbeige4-3B-Thinking (2511)Nanbeige
173.0%
–
375 gpt-oss-20b (medium reasoning)OpenAI
186.0%
$0.03 / $0.15
376 Magistral Small 1.2 (reasoning by system prompt)Mistral AI
197.0%
–
377 Ministral 3 8B Reasoning (reasoning by system prompt)Mistral AI
224.0%
–
378 Nemotron 3 Nano 30B A3B (reasoning by prefill)NVIDIA
251.0%
$0.05 / $0.20
379 gpt-oss-20b (high reasoning)OpenAI
282.0%
$0.03 / $0.15

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Results as published by UGI Leaderboard; we do not re-run them.

What it measures

Average percentage difference between a requested and the delivered word count.

What it does not measure

Not other format limits such as character counts or bullet counts.

947 results from UGI Leaderboard not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Community fine-tunes and merges: 934
  • No score published: 13

UGI Leaderboard by DontPlanToEnd, Apache License 2.0.