Benchmarks / UGI Leaderboard

Reported by UGI Leaderboard

UGI Leaderboard

How closely the model matches the writing style of a given example.

Last updated 2 Oct 2026

Results dated
6 Sep 2025 to 2 Oct 2026
Results
379 configurations of 230 models
Unit
score from 0 to 1
Licence
Apache License 2.0

Style adherence: Trinity Large Preview

Top 15 of 230 results · score from 0 to 1, higher is better · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ Claude Opus 5.5 (max reasoning)Anthropic 0.43
  2. 1≈ Claude Fable 5.1 (high reasoning)Anthropic 0.43
  3. 3 Claude Opus 4.6 (high reasoning)Anthropic 0.42
  4. 4 Claude Opus 4.7 (high reasoning)Anthropic 0.41
  5. 4 Claude Sonnet 4.6 (max reasoning)Anthropic 0.41
  6. 4 GPT-5.5 (high reasoning)OpenAI 0.41
  7. 4 MiniMax-M2.1MiniMax 0.41
  8. 4 Claude Sonnet 4.5 (reasoning on)Anthropic 0.41
  9. 4 Claude Fable 5 (extra-high reasoning)Anthropic 0.41
  10. 10 Claude Opus 4.8 (max reasoning)Anthropic 0.40
  11. 10 MiMo-V2.5-Pro (reasoning on)Xiaomi 0.40
  12. 10 MiMo-V2.5 (reasoning on)Xiaomi 0.40
  13. 10 GLM-5 (no reasoning)Z.ai 0.40
  14. 10 Trinity Large PreviewArcee AI 0.40
  15. 10 GPT-5.4 (high reasoning)OpenAI 0.40

Full results

UGI: style adherence, score from 0 to 1, higher is better
#ModelStyle adherence
score from 0 to 1, higher is better
Price
$ per million tokens, in / out
1≈ Claude Opus 5.5 (max reasoning)Anthropic · best of 5 settings
0.43
$4 / $20
1≈ Claude Fable 5.1 (high reasoning)Anthropic · best of 3 settings
0.43
$10 / $50
3 Claude Opus 4.6 (high reasoning)Anthropic · best of 4 settings
0.42
$5 / $25
4 Claude Opus 4.7 (high reasoning)Anthropic · best of 4 settings
0.41
$5 / $25
4 Claude Sonnet 4.6 (max reasoning)Anthropic · best of 4 settings
0.41
$3 / $15
4 GPT-5.5 (high reasoning)OpenAI · best of 5 settings
0.41
$5 / $30
4 MiniMax-M2.1MiniMax
0.41
$0.30 / $1.20
4 Claude Sonnet 4.5 (reasoning on)Anthropic · best of 2 settings
0.41
$3 / $15
4 Claude Fable 5 (extra-high reasoning)Anthropic · best of 5 settings
0.41
$10 / $50
10 Claude Opus 4.8 (max reasoning)Anthropic · best of 5 settings
0.40
$5 / $25
10 MiMo-V2.5-Pro (reasoning on)Xiaomi · best of 2 settings
0.40
$0.43 / $0.87
10 MiMo-V2.5 (reasoning on)Xiaomi · best of 2 settings
0.40
$0.17 / $0.34
10 GLM-5 (no reasoning)Z.ai · best of 2 settings
0.40
$0.95 / $2.55
10 Trinity Large PreviewArcee AI
0.40
–
10 GPT-5.4 (high reasoning)OpenAI · best of 5 settings
0.40
$2.50 / $15
10 Gemini 2.5 ProGoogle
0.40
$1.25 / $10
10 GPT-5.2 (low reasoning)OpenAI · best of 4 settings
0.40
$1.75 / $14
18 GPT-5.6 Terra (extra-high reasoning)OpenAI · best of 5 settings
0.39
$2 / $12
18 GPT-6 Sol (extra-high reasoning)OpenAI · best of 4 settings
0.39
$2 / $10
18 Claude Opus 4.5 (reasoning on)Anthropic · best of 2 settings
0.39
$5 / $25
18 GPT-5 (high reasoning)OpenAI · best of 2 settings
0.39
$1.25 / $10
18 Claude Haiku 4.5 (reasoning on)Anthropic · best of 2 settings
0.39
$1 / $5
18 GLM-4.6V (no reasoning)Z.ai · best of 2 settings
0.39
$0.30 / $0.90
18 Mistral Medium 3.1Mistral AI
0.39
$0.40 / $2
25 Kimi K2 ThinkingMoonshot AI
0.38
$0.60 / $2.50
25 MiniMax-M2.7MiniMax
0.38
$0.30 / $1.20
25 GPT-5.1 (low reasoning)OpenAI · best of 3 settings
0.38
$1.25 / $10
25 GLM-5.1 (reasoning on)Z.ai · best of 2 settings
0.38
$1.38 / $4.40
25 Qwen3-1.7B (no reasoning)Alibaba · best of 2 settings
0.38
–
25 GPT-5.6 Luna (extra-high reasoning)OpenAI · best of 5 settings
0.38
$0.20 / $1.20
25 Claude Sonnet 5 (extra-high reasoning)Anthropic · best of 5 settings
0.38
$2 / $10
25 Llama 3.1 70B InstructMeta
0.38
$0.40 / $0.40
25 GLM-4.6 (reasoning on)Z.ai · best of 2 settings
0.38
$0.50 / $2
25 GLM-4.7 (reasoning on)Z.ai · best of 2 settings
0.38
$0.54 / $1.98
25 MiniMax-M2.5MiniMax
0.38
$0.30 / $1.20
25 Mistral Large 3 (no reasoning)Mistral AI · best of 2 settings
0.38
$0.50 / $1.50
25 GPT-6 Astra (medium reasoning)OpenAI · best of 4 settings
0.38
$10 / $50
25 Claude Opus 4.1 (reasoning on)Anthropic · best of 2 settings
0.38
$15 / $75
39 GPT-6.1 Sol (high reasoning)OpenAI · best of 4 settings
0.37
$2 / $10
39 Step 3.5 FlashStepFun
0.37
$0.10 / $0.30
39 GLM-5.2 (no reasoning)Z.ai · best of 2 settings
0.37
$1.40 / $4.40
39 DeepSeek-V3DeepSeek
0.37
$0.26 / $1.03
39 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek · best of 2 settings
0.37
$0.27 / $1
39 GPT-5.6 Sol (extra-high reasoning)OpenAI · best of 5 settings
0.37
$4 / $20
39 Qwen3-4B (no reasoning)Alibaba · best of 2 settings
0.37
–
39 Llama 4 MaverickMeta
0.37
$0.27 / $0.85
39 Muse Glimmer 30B (medium reasoning)Meta · best of 4 settings
0.37
$0.30 / $1.20
39 GLM-4.7-Flash (reasoning by prefill)Z.ai · best of 2 settings
0.37
$0.06 / $0.40
39 Qwen3-14B (no reasoning)Alibaba · best of 2 settings
0.37
$0.12 / $0.24
39 Gemini 3 Pro Preview (high reasoning)Google · best of 2 settings
0.37
–
39 Llama 3.3 70B InstructMeta
0.37
$0.59 / $0.79
39 Gemini 3 Flash Preview (high reasoning)Google · best of 3 settings
0.37
$0.50 / $3
39 DeepSeek-V3.2-SpecialeDeepSeek
0.37
–
39 Grok 4.7xAI
0.37
$2 / $6
39 Jamba Mini 1.7AI21 Labs
0.37
–
39 DeepSeek-V3.2 (reasoning on)DeepSeek · best of 2 settings
0.37
$0.30 / $0.96
39 DeepSeek-V4-Flash (0423, reasoning on)DeepSeek · best of 2 settings
0.37
$0.14 / $0.28
39 MiniMax-M2MiniMax
0.37
$0.30 / $1.20
39 Mistral Large 2.1 (2411)Mistral AI
0.37
–
60 Qwen3-8B (no reasoning)Alibaba · best of 2 settings
0.36
$0.12 / $0.46
60 Claude Opus 4 (reasoning on)Anthropic · best of 2 settings
0.36
–
60 Llama 3.1 405B InstructMeta
0.36
–
60 Mistral Large 2 (2407)Mistral AI
0.36
$2 / $6
60 Qwen2.5-72B-InstructAlibaba
0.36
$0.36 / $0.40
60 Qwen3-235B-A22B-Instruct-2507Alibaba
0.36
$0.15 / $0.75
60 Qwen3-30B-A3B (no reasoning)Alibaba · best of 2 settings
0.36
$0.12 / $0.50
60 DeepSeek-V3.1 (reasoning on)DeepSeek · best of 2 settings
0.36
$0.55 / $1.65
60 Llama 4 ScoutMeta
0.36
$0.18 / $0.59
60 GPT-5.3 ChatOpenAI
0.36
–
60 o1 (high reasoning)OpenAI · best of 2 settings
0.36
$15 / $60
60 Grok 3xAI
0.36
–
60 MiMo-V2-Flash (no reasoning)Xiaomi · best of 2 settings
0.36
–
60 Gemma 4 31B (reasoning by prefill)Google · best of 2 settings
0.36
$0.14 / $0.40
60 Falcon-H1 0.5B InstructTII
0.36
–
60 GLM-4.5 (reasoning on)Z.ai · best of 2 settings
0.36
$0.60 / $2.20
60 Jamba Large 1.7AI21 Labs
0.36
–
60 Claude Sonnet 4 (no reasoning)Anthropic · best of 2 settings
0.36
$3 / $15
60 DeepSeek-V4-Pro (0423, reasoning on)DeepSeek · best of 2 settings
0.36
$1.42 / $2.83
60 Rnj-1 InstructEssential AI
0.36
–
60 Mixtral 8x7B Instruct v0.1Mistral AI
0.36
–
60 Grok 4.6xAI
0.36
$2 / $6
60 Llama 3.2 3B InstructMeta
0.36
$0.05 / $0.33
60 Qwen3.5-122B-A10B (no reasoning)Alibaba · best of 2 settings
0.36
$0.26 / $2.08
60 Grok 4 (0709)xAI
0.36
–
60 Grok 4.20 Multi-Agent Beta (0309, 4 agents)xAI
0.36
–
60 Qwen3-VL-235B-A22B-InstructAlibaba
0.36
$0.30 / $1.50
60 Seed-OSS-36B-InstructByteDance · best of 3 settings
0.36
–
88 Claude 3.7 Sonnet (reasoning on)Anthropic · best of 2 settings
0.35
–
88 Llama 2 70B ChatMeta
0.35
–
88 Mistral Small 3.1 24BMistral AI
0.35
$0.35 / $0.56
88 ChatGPT-4o (2025-03-26)OpenAI
0.35
–
88 Grok 4.20 Beta (0309, non-reasoning)xAI
0.35
–
88 Command R (08-2024)Cohere
0.35
$0.15 / $0.60
88 Gemma 4 26B A4B (reasoning by prefill)Google · best of 2 settings
0.35
$0.10 / $0.30
88 Mistral Medium 3Mistral AI
0.35
$0.40 / $2
88 Qwen3.6-Plus (no reasoning)Alibaba · best of 2 settings
0.35
$0.33 / $1.95
88 Gemini 3.5 Flash (medium reasoning)Google · best of 4 settings
0.35
$1.50 / $9
88 Granite 4.0 H SmallIBM
0.35
–
88 GPT-5 Chat (latest)OpenAI
0.35
–
88 Qwen3-VL-32B-InstructAlibaba
0.35
$0.10 / $0.42
88 Llama 3.3 8B InstructAllura Forge
0.35
–
88 Mixtral 8x22B InstructMistral AI
0.35
$2 / $6
88 Mistral Small 3Mistral AI
0.35
$0.05 / $0.08
88 DeepSeek-V3-0324DeepSeek
0.35
$0.25 / $1
88 Mistral Small 3.2 24BMistral AI
0.35
$0.094 / $0.25
88 Mistral Small (2409)Mistral AI
0.35
–
88 GPT-5.2 ChatOpenAI
0.35
–
88 Qwen3.6-27B (no reasoning)Alibaba · best of 2 settings
0.35
$0.30 / $3.20
88 Nova 2 Lite (no reasoning)Amazon · best of 2 settings
0.35
$0.30 / $2.50
88 Ministral 3 14B ReasoningMistral AI · best of 2 settings
0.35
–
88 Mistral Medium 3.5 (no reasoning)Mistral AI · best of 2 settings
0.35
$1.50 / $7.50
88 Kimi K2.6 (reasoning on)Moonshot AI · best of 2 settings
0.35
$0.95 / $4
88 Claude 3 OpusAnthropic
0.35
–
88 Command ACohere
0.35
$2.50 / $10
88 DeepSeek-V3.2-Exp (no reasoning)DeepSeek · best of 2 settings
0.35
$0.27 / $0.41
88 Devstral Small 2Mistral AI
0.35
–
88 Qwen3-4B-Instruct-2507Alibaba
0.35
–
88 Gemini 2.5 Flash Preview (09-2025, reasoning on)Google · best of 2 settings
0.35
–
88 Gemma 4 12BGoogle · best of 2 settings
0.35
–
88 Ministral 3 8B Reasoning (reasoning by prefill)Mistral AI · best of 3 settings
0.35
–
88 Hy3 Preview (no reasoning)Tencent · best of 2 settings
0.35
$0.18 / $0.60
122 Qwen3-VL-235B-A22B-ThinkingAlibaba
0.34
$0.40 / $4
122 Gemini 3.1 Pro Preview (low reasoning)Google · best of 3 settings
0.34
$2 / $12
122 Llama 3.1 8B InstructMeta
0.34
$0.05 / $0.08
122 Phi-4Microsoft
0.34
$0.07 / $0.14
122 Inflection 3 PiInflection AI
0.34
–
122 Magistral Small 1.2Mistral AI · best of 2 settings
0.34
–
122 Kimi K2.5 (no reasoning)Moonshot AI · best of 2 settings
0.34
$0.57 / $2.85
122 Grok 4.20 (0309, non-reasoning)xAI
0.34
–
122 Command R+ (08-2024)Cohere
0.34
$2.50 / $10
122 Claude 3 HaikuAnthropic
0.34
–
122 Gemini 3.8 Flash (high reasoning)Google · best of 3 settings
0.34
$1.50 / $7.50
122 Kimi-VL-A3B-InstructMoonshot AI
0.34
–
122 EXAONE 4.0 32BLG AI Research
0.34
–
122 Mistral NemoMistral AI
0.34
$0.023 / $0.03
122 GPT-5.1 ChatOpenAI
0.34
–
122 Grok 4.3xAI
0.34
$1.25 / $2.50
122 Qwen3-VL-32B-ThinkingAlibaba
0.34
–
139 Kimi-VL-A3B-Thinking (2506)Moonshot AI
0.33
–
139 Qwen3-30B-A3B-Thinking-2507Alibaba
0.33
$0.20 / $2.40
139 Qwen3-VL-4B-ThinkingAlibaba
0.33
–
139 Qwen3-VL-8B-InstructAlibaba
0.33
$0.12 / $0.46
139 MedGemma 27B TextGoogle
0.33
–
139 Llama 3.2 1B InstructMeta
0.33
$0.027 / $0.20
139 Ministral 3 8B (reasoning by prefill)Mistral AI · best of 2 settings
0.33
$0.15 / $0.15
139 o4-mini (low reasoning)OpenAI · best of 3 settings
0.33
$1.10 / $4.40
139 Qwen3.5-27B (no reasoning)Alibaba · best of 2 settings
0.33
$0.27 / $2.16
139 Qwen3-VL-2B-ThinkingAlibaba
0.33
–
139 Grok 4.1 Fast (reasoning)xAI
0.33
–
139 Grok 4.20 Beta (0309, reasoning)xAI
0.33
–
139 Qwen3.5-2B (reasoning by prefill)Alibaba · best of 2 settings
0.33
–
139 Qwen3.5-35B-A3B (no reasoning)Alibaba · best of 2 settings
0.33
$0.16 / $1.30
139 GPT-4o (2024-05-13)OpenAI
0.33
$5 / $15
139 Ling-1TAnt Group
0.33
–
139 Qwen3-MaxAlibaba
0.33
$0.78 / $3.90
139 Qwen3-VL-8B-ThinkingAlibaba
0.33
$0.18 / $2.10
139 Grok 4 Fast (reasoning)xAI
0.33
–
139 GLM-4.5-AirZ.ai · best of 2 settings
0.33
$0.14 / $0.86
139 Gemini 3.6 Flash (medium reasoning)Google · best of 3 settings
0.33
$1.50 / $7.50
139 Gemini 3.7 Flash (medium reasoning)Google · best of 3 settings
0.33
$1.50 / $7.50
139 GPT-4.1OpenAI
0.33
$2 / $8
139 Falcon-H1 3B InstructTII
0.33
–
139 Grok 4 Fast (non-reasoning)xAI
0.33
–
139 Qwen3.5-397B-A17B (no reasoning)Alibaba
0.33
$0.55 / $3.50
139 Falcon-H1 7B InstructTII
0.33
–
139 Grok 4.20 (0309, reasoning)xAI
0.33
–
167 Ministral 3 14BMistral AI · best of 2 settings
0.32
$0.20 / $0.20
167 Gemma 4 E2B (reasoning by prefill)Google · best of 2 settings
0.32
–
167 Qwen2.5-VL-72B-InstructAlibaba
0.32
$0.80 / $1
167 Falcon-H1 1.5B InstructTII
0.32
–
167 Solar Pro 3Upstage
0.32
$0.15 / $0.60
167 GLM-4-32B-0414Z.ai
0.32
–
167 Qwen3.5-9B (no reasoning)Alibaba · best of 2 settings
0.32
$0.10 / $0.15
167 Qwen3.6-35B-A3B (no reasoning)Alibaba · best of 2 settings
0.32
$0.10 / $1
167 Qwen3-32B (no reasoning)Alibaba · best of 2 settings
0.32
$0.14 / $0.40
167 Qwen3.5-4B (no reasoning)Alibaba · best of 2 settings
0.32
–
167 Kimi Linear 48B A3B InstructMoonshot AI
0.32
–
167 Qwen2.5-7B-InstructAlibaba
0.32
$0.10 / $0.20
167 Qwen2.5-VL-32B-InstructAlibaba
0.32
–
167 Ring-1TAnt Group
0.32
–
167 Grok 4.1 Fast (non-reasoning)xAI
0.32
–
167 Grok 4.5xAI
0.32
$2 / $6
167 Qwen2.5-Coder-7B-InstructAlibaba
0.32
–
167 Qwen3.5-0.8B (no reasoning)Alibaba
0.32
–
167 Olmo 3 32B ThinkAi2
0.32
–
167 DeepSeek-R1-0528DeepSeek
0.32
$0.50 / $2.18
187 Qwen3-30B-A3B-Instruct-2507Alibaba
0.31
$0.09 / $0.30
187 gpt-oss-20b (medium reasoning)OpenAI · best of 3 settings
0.31
$0.03 / $0.15
187 Qwen3-VL-30B-A3B-ThinkingAlibaba
0.31
$0.29 / $1
187 Falcon-H1 1.5B Deep InstructTII
0.31
–
187 Apriel Nemotron 15B ThinkerServiceNow
0.31
–
187 Mistral Small 4 (high reasoning)Mistral AI · best of 2 settings
0.31
$0.15 / $0.60
187 Reka Flash 3Reka AI
0.31
$0.10 / $0.20
187 Qwen3-235B-A22B-Thinking-2507Alibaba
0.31
$0.30 / $3
187 Kimi K2 (0905)Moonshot AI
0.31
$0.60 / $2.50
187 Qwen3-Coder-30B-A3B-InstructAlibaba
0.31
$0.07 / $0.28
187 Qwen3-Next-80B-A3B-InstructAlibaba
0.31
$0.10 / $1.10
187 o3 (high reasoning)OpenAI · best of 3 settings
0.31
$2 / $8
199 Qwen2.5-32B-InstructAlibaba
0.30
–
199 Gemma 2 27BGoogle
0.30
$0.65 / $0.65
199 Gemma 4 E4B (reasoning by prefill)Google · best of 2 settings
0.30
–
199 QwQ-32BAlibaba
0.30
–
199 Qwen2.5-VL-3B-InstructAlibaba
0.30
–
199 DeepSeek-R1DeepSeek
0.30
$0.70 / $2.50
199 InternLM3 8B InstructShanghai AI Lab
0.30
–
199 Qwen2.5-1.5B-InstructAlibaba
0.30
–
199 Inflection 3 ProductivityInflection AI
0.30
–
199 Gemma 2 2BGoogle
0.30
–
199 Gemma 2 9BGoogle
0.30
–
199 Olmo 3 7B InstructAi2
0.30
–
211 Qwen3-Omni-30B-A3B-ThinkingAlibaba
0.29
–
211 Gemma 3 12BGoogle
0.29
$0.05 / $0.15
211 Gemma 3 27BGoogle
0.29
$0.12 / $0.20
211 Gemma 3 4BGoogle
0.29
$0.05 / $0.10
211 Olmo 3 7B ThinkAi2
0.29
–
211 Qwen3-VL-4B-InstructAlibaba
0.29
–
211 Qwen2.5-14B-InstructAlibaba
0.29
–
211 Kimi K2 (0711)Moonshot AI
0.29
$0.57 / $2.30
211 gpt-oss-120b (medium reasoning)OpenAI
0.29
$0.15 / $0.60
211 Qwen3-Next-80B-A3B-ThinkingAlibaba
0.29
$0.15 / $1.20
221 Nemotron 3 Nano 30B A3BNVIDIA · best of 2 settings
0.28
$0.05 / $0.20
221 Nanbeige4-3B-Thinking (2510)Nanbeige
0.28
–
221 Solar 10.7B Instruct v1.0Upstage
0.28
–
221 Qwen3-4B-Thinking-2507Alibaba
0.28
–
221 Qwen3-VL-2B-InstructAlibaba
0.28
–
221 LFM2-8B-A1BLiquid AI
0.28
–
227 LFM2-24B-A2BLiquid AI
0.27
–
227 Nemotron Nano 12B v2 VL (BF16, no reasoning)NVIDIA · best of 2 settings
0.27
–
229 Nanbeige4-3B-Thinking (2511)Nanbeige
0.26
–
230 Qwen3-0.6B (reasoning on)Alibaba
0.25
–

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Results as published by UGI Leaderboard; we do not re-run them.

What it measures

How closely the model matches the writing style of a given example.

What it does not measure

Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.

947 results from UGI Leaderboard not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Community fine-tunes and merges: 934
  • No score published: 13

UGI Leaderboard by DontPlanToEnd, Apache License 2.0.