Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Human preference on creative writing prompts.

Last updated 2 Oct 2026

Results dated
2 Oct 2026
Results
411 configurations of 385 models
Unit
Arena rating
Licence
Creative Commons Attribution 4.0 International

Creative writing: Claude Opus 4.8

Top 15 of 385 results · Arena rating, higher is better · lines show the 95% range · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ Gemini 4 Argon (high reasoning)Google 1,519
  2. 2≈ Claude Opus 5.5 (high reasoning)Anthropic 1,516
  3. 3≈ Claude Fable 5 (high reasoning)Anthropic 1,503
  4. 4≈ Claude Opus 4.6 (high reasoning)Anthropic 1,501
  5. 5≈ Gemini 3.7 Flash (high reasoning)Google 1,493
  6. 6 Claude Opus 4.7 (high reasoning)Anthropic 1,489
  7. 7 Gemini 3.8 Flash (high reasoning)Google 1,485
  8. 8 Gemini 3 Pro PreviewGoogle 1,484
  9. 9 Claude Fable 5.1 (max reasoning)Anthropic 1,482
  10. 10 Gemini 3.1 Pro PreviewGoogle 1,480
  11. 11 Claude Opus 5 (high reasoning)Anthropic 1,471
  12. 11 Gemini 3.6 Flash (high reasoning)Google 1,471
  13. 13 Claude Opus 4.5 (high reasoning, 32k budget)Anthropic 1,470
  14. 13 Claude Opus 4.8 (high reasoning)Anthropic 1,470
  15. 15 Gemini 3.5 Flash (medium reasoning)Google 1,469

Full results

Arena Text: creative writing, Arena rating, higher is better
#ModelCreative writing · 95% range
Arena rating, higher is better
Price
$ per million tokens, in / out
1≈ Gemini 4 Argon (high reasoning)Google
1,519
1,499–1,538
–
2≈ Claude Opus 5.5 (high reasoning)Anthropic
1,516
1,497–1,536
$4 / $20
3≈ Claude Fable 5 (high reasoning)Anthropic
1,503
1,495–1,510
$10 / $50
4≈ Claude Opus 4.6 (high reasoning)Anthropic · best of 2 settings
1,501
1,494–1,507
$5 / $25
5≈ Gemini 3.7 Flash (high reasoning)Google
1,493
1,484–1,503
$1.50 / $7.50
6 Claude Opus 4.7 (high reasoning)Anthropic · best of 2 settings
1,489
1,482–1,496
$5 / $25
7 Gemini 3.8 Flash (high reasoning)Google
1,485
1,476–1,494
$1.50 / $7.50
8 Gemini 3 Pro PreviewGoogle
1,484
1,476–1,493
–
9 Claude Fable 5.1 (max reasoning)Anthropic
1,482
1,469–1,494
$10 / $50
10 Gemini 3.1 Pro PreviewGoogle
1,480
1,475–1,486
$2 / $12
11 Claude Opus 5 (high reasoning)Anthropic · best of 2 settings
1,471
1,465–1,478
$5 / $25
11 Gemini 3.6 Flash (high reasoning)Google
1,471
1,463–1,479
$1.50 / $7.50
13 Claude Opus 4.5 (high reasoning, 32k budget)Anthropic · best of 2 settings
1,470
1,461–1,478
$5 / $25
13 Claude Opus 4.8 (high reasoning)Anthropic · best of 2 settings
1,470
1,463–1,476
$5 / $25
15 Gemini 3.5 Flash (medium reasoning)Google · best of 2 settings
1,469
1,461–1,476
$1.50 / $9
16 GPT-5.6 Sol (extra-high reasoning)OpenAI
1,468
1,460–1,476
$4 / $20
16 Qwen3.8-Max (0902)Alibaba
1,468
1,458–1,478
$2 / $6
18 Muse SparkMeta
1,465
1,452–1,479
–
19 Grok 4.20 Beta 1xAI
1,463
1,453–1,473
–
20 Kimi K3 (max reasoning)Moonshot AI
1,460
1,451–1,469
$3 / $15
20 GPT-6.1 Sol (max reasoning)OpenAI
1,460
1,436–1,483
$2 / $10
22 Muse Spark 1.3 (max reasoning)Meta
1,459
1,446–1,471
$1.25 / $4.25
23 Gemini 3 Flash PreviewGoogle · best of 2 settings
1,457
1,448–1,466
$0.50 / $3
24 GPT-5.5 InstantOpenAI
1,456
1,447–1,466
–
25 Muse Spark 1.2 (extra-high reasoning)Meta
1,455
1,435–1,476
$1.25 / $4.25
26 GLM-5.2 (max reasoning)Z.ai
1,454
1,447–1,462
$1.40 / $4.40
26 Claude Sonnet 4.5Anthropic · best of 2 settings
1,454
1,448–1,460
$3 / $15
26 GLM-5.3 (max reasoning)Z.ai
1,454
1,444–1,464
$1.40 / $4.40
29 GPT-5.5 (high reasoning)OpenAI · best of 2 settings
1,451
1,445–1,458
$5 / $30
30 Grok 4.5xAI
1,450
1,443–1,458
$2 / $6
31 Claude Sonnet 4.6Anthropic
1,449
1,443–1,456
$3 / $15
31 GPT-6 Astra (max reasoning)OpenAI
1,449
1,435–1,463
$10 / $50
31 GLM-5Z.ai
1,449
1,440–1,458
$0.95 / $2.55
34 Claude Sonnet 5.5 (extra-high reasoning)Anthropic
1,447
1,424–1,471
$2 / $10
34 Muse Spark 1.1Meta
1,447
1,439–1,455
$1.25 / $4.25
34 Grok 4.20 Multi-Agent Beta (0309)xAI
1,447
1,440–1,454
–
37 GLM-5.1Z.ai
1,446
1,440–1,453
$1.38 / $4.40
37 MiMo-V2.6-ProXiaomi
1,446
1,425–1,467
$0.43 / $0.87
37 DeepSeek-V4-Pro (0423)DeepSeek · best of 2 settings
1,446
1,439–1,453
$1.42 / $2.83
37 Claude Opus 4.1 (16k reasoning budget)Anthropic · best of 2 settings
1,446
1,438–1,453
$15 / $75
41 Grok 4.20 Beta (0309, reasoning)xAI
1,445
1,439–1,452
–
41 Qwen3.5-Max-PreviewAlibaba
1,445
1,434–1,456
–
41 Grok 4.6 (high reasoning)xAI
1,445
1,436–1,454
$2 / $6
41 DeepSeek-V4-Pro (0813, high reasoning)DeepSeek
1,445
1,431–1,458
$1.32 / $3.96
45 Gemini 2.5 ProGoogle
1,443
1,438–1,448
$1.25 / $10
46 GPT-5.4 (high reasoning)OpenAI · best of 2 settings
1,442
1,436–1,449
$2.50 / $15
46 GPT-6 Sol (max reasoning)OpenAI
1,442
1,427–1,456
$2 / $10
48 Qwen3.7-Max-PreviewAlibaba
1,441
1,414–1,469
–
49 DeepSeek-V4.1-Flash (max reasoning)DeepSeek
1,440
1,425–1,455
$0.30 / $1.20
50 Qwen3.6-Max-PreviewAlibaba
1,437
1,414–1,459
$1.03 / $6.16
51 GPT-4.5 PreviewOpenAI
1,436
1,425–1,448
–
51 Gemini 3.5 Flash-LiteGoogle
1,436
1,428–1,444
$0.30 / $2.50
51 Grok 4.7 (extra-high reasoning)xAI
1,436
1,419–1,452
$2 / $6
54 Claude Sonnet 5 (high reasoning)Anthropic
1,435
1,428–1,442
$2 / $10
55 MiMo-V2.5-ProXiaomi
1,434
1,428–1,440
$0.43 / $0.87
55 GPT-5.2 Chat (2026-02-10)OpenAI
1,434
1,425–1,442
–
57 GLM-5.3-FlashZ.ai
1,433
1,424–1,443
$0.15 / $0.50
57 Kimi K2.6Moonshot AI
1,433
1,425–1,441
$0.95 / $4
59 Claude Opus 4 (16k reasoning budget)Anthropic · best of 2 settings
1,432
1,423–1,442
–
59 Qwen3.7-PlusAlibaba
1,432
1,424–1,439
$0.32 / $1.28
61 Grok 4.1xAI
1,431
1,425–1,438
–
62 GPT-5.1 (high reasoning)OpenAI · best of 2 settings
1,428
1,420–1,437
$1.25 / $10
62 ERNIE 5.1Baidu
1,428
1,419–1,436
–
64 Grok 4.1 ThinkingxAI
1,426
1,420–1,433
–
64 GPT-5.6 Terra (extra-high reasoning)OpenAI
1,426
1,418–1,433
$2 / $12
66 MiMo-V2-ProXiaomi
1,424
1,413–1,434
–
67 Kimi K2.5 (reasoning on)Moonshot AI · best of 2 settings
1,423
1,417–1,430
$0.57 / $2.85
68 Grok 4.3xAI
1,422
1,416–1,429
$1.25 / $2.50
68 Hy3Tencent
1,422
1,408–1,435
$0.14 / $0.58
68 ChatGPT-4o (2025-03-26)OpenAI
1,422
1,416–1,428
–
68 ERNIE 5.0 (0110)Baidu
1,422
1,413–1,430
–
72 Step 5 PreviewStepFun
1,421
1,396–1,445
–
73 ERNIE 5.0 Preview (1203)Baidu
1,420
1,405–1,435
–
73 Gemma 4 31BGoogle
1,420
1,401–1,439
$0.14 / $0.40
75 Grok 4.1 Fast (reasoning)xAI
1,414
1,408–1,421
–
76 Gemini 3.1 Flash-Lite PreviewGoogle
1,413
1,406–1,420
$0.25 / $1.50
76 GPT-6 Luna (max reasoning)OpenAI
1,413
1,400–1,427
$0.10 / $0.50
78 DeepSeek-V4-Flash (0423, high reasoning)DeepSeek · best of 2 settings
1,410
1,403–1,418
$0.14 / $0.28
78 DeepSeek-V3.2-ExpDeepSeek · best of 2 settings
1,410
1,395–1,424
$0.27 / $0.41
80 GPT-5.6 Luna (extra-high reasoning)OpenAI
1,409
1,401–1,417
$0.20 / $1.20
81 GPT-5.3 ChatOpenAI
1,408
1,399–1,417
–
82 Qwen3.6-PlusAlibaba
1,407
1,399–1,415
$0.33 / $1.95
82 ERNIE 5.0 Preview (1022)Baidu
1,407
1,385–1,429
–
84 MiniMax-M3MiniMax
1,406
1,399–1,413
$0.30 / $1.20
84 GLM-5V-TurboZ.ai
1,406
1,391–1,420
$1.20 / $4
86 Qwen3.5-397B-A17BAlibaba
1,405
1,400–1,411
$0.55 / $3.50
86 DeepSeek-V3.1-TerminusDeepSeek · best of 2 settings
1,405
1,378–1,433
$0.27 / $1
86 GLM-4.7Z.ai
1,405
1,391–1,418
$0.54 / $1.98
89 Dola-Seed-2.0-ProByteDance
1,403
1,397–1,410
–
89 GPT-4.1OpenAI
1,403
1,395–1,410
$2 / $8
91 DeepSeek-V3.1 (reasoning on)DeepSeek · best of 2 settings
1,402
1,387–1,417
$0.55 / $1.65
92 GPT-5.4 mini (high reasoning)OpenAI
1,401
1,394–1,409
$0.75 / $4.50
92 DeepSeek-V3.2DeepSeek · best of 2 settings
1,401
1,393–1,408
$0.30 / $0.96
94 GLM-4.6Z.ai
1,400
1,392–1,409
$0.50 / $2
94 Mistral Medium 3.5Mistral AI
1,400
1,386–1,414
$1.50 / $7.50
94 Gemma 4 26B A4BGoogle
1,400
1,381–1,418
$0.10 / $0.30
97 Grok 3 Preview (02-24)xAI
1,398
1,389–1,407
–
98 Claude Sonnet 4 (32k reasoning budget)Anthropic · best of 2 settings
1,396
1,386–1,405
$3 / $15
99 MiMo-V2-OmniXiaomi
1,395
1,384–1,406
–
99 Grok 4 (0709)xAI
1,395
1,386–1,403
–
101 MiMo-V2.5Xiaomi
1,394
1,387–1,402
$0.17 / $0.34
101 Gemini 2.5 FlashGoogle
1,394
1,389–1,399
$0.30 / $2.50
103 Qwen3-Max-PreviewAlibaba
1,393
1,383–1,404
–
103 Claude 3.7 Sonnet (32k reasoning budget)Anthropic · best of 2 settings
1,393
1,385–1,402
–
103 Qwen3-MaxAlibaba
1,393
1,376–1,410
$0.78 / $3.90
106 GPT-5.2OpenAI · best of 2 settings
1,391
1,385–1,398
$1.75 / $14
106 Kimi K2 Thinking TurboMoonshot AI
1,391
1,384–1,398
–
108 DeepSeek-V3-0324DeepSeek
1,390
1,382–1,398
$0.25 / $1
108 DeepSeek-R1-0528DeepSeek
1,390
1,377–1,402
$0.50 / $2.18
108 LongCat-Flash-Chat (2602, experimental)Meituan
1,390
1,380–1,399
–
111 GPT-5 ChatOpenAI
1,389
1,380–1,399
–
112 Claude Haiku 4.5Anthropic
1,388
1,383–1,393
$1 / $5
113 MiMo-V2.6-FlashXiaomi
1,385
1,367–1,404
$0.14 / $0.28
114 Grok 4 Fast ChatxAI
1,384
1,364–1,404
–
115 Grok 4 Fast (reasoning)xAI
1,383
1,371–1,395
–
115 o3OpenAI
1,383
1,376–1,390
$2 / $8
117 Nemotron 3 Ultra 550B A55B (NVFP4)NVIDIA
1,382
1,368–1,395
–
118 o1OpenAI
1,381
1,373–1,390
$15 / $60
118 Kimi K2 (0905)Moonshot AI
1,381
1,366–1,397
$0.60 / $2.50
118 Gemini 2.5 Flash Preview (09-2025)Google
1,381
1,372–1,390
–
121 InklingThinking Machines
1,379
1,371–1,387
$0.95 / $4.05
122 Qwen3-235B-A22B-Instruct-2507Alibaba
1,378
1,373–1,384
$0.15 / $0.75
122 Mistral Medium 3.1Mistral AI
1,378
1,373–1,384
$0.40 / $2
122 Hunyuan Vision 1.5 ThinkingTencent
1,378
1,342–1,415
–
125 DeepSeek-R1DeepSeek
1,374
1,364–1,384
$0.70 / $2.50
125 Mistral Large 3Mistral AI
1,374
1,368–1,379
$0.50 / $1.50
127 Hunyuan T1 (2025-07-11)Tencent
1,373
1,350–1,397
–
127 GPT-5 (high reasoning)OpenAI
1,373
1,363–1,383
$1.25 / $10
129 Gemini 2.5 Flash-Lite Preview (06-17, reasoning on)Google
1,372
1,363–1,381
–
129 GLM-4.5Z.ai
1,372
1,361–1,383
$0.60 / $2.20
131 Kimi K2 (0711)Moonshot AI
1,371
1,361–1,381
$0.57 / $2.30
131 Claude 3.5 Sonnet (2024-10-22)Anthropic
1,371
1,365–1,377
–
131 Qwen3-235B-A22B-Thinking-2507Alibaba
1,371
1,352–1,389
$0.30 / $3
134 Qwen3-235B-A22B (no reasoning)Alibaba · best of 2 settings
1,366
1,357–1,375
$0.46 / $1.82
134 o1-previewOpenAI
1,366
1,356–1,376
–
134 Qwen3.5-122B-A10BAlibaba
1,366
1,356–1,375
$0.26 / $2.08
134 Muse GlimmerMeta
1,366
1,343–1,388
–
138 Mistral Medium 3Mistral AI
1,364
1,355–1,374
$0.40 / $2
138 MiniMax-M2.7MiniMax
1,364
1,358–1,370
$0.30 / $1.20
138 Qwen3.8-27BAlibaba
1,364
1,353–1,374
$0.50 / $3
141 Qwen3-Coder-480B-A35BAlibaba
1,363
1,352–1,373
$0.35 / $1.50
142 Hunyuan TurboS (2025-04-16)Tencent
1,361
1,346–1,377
–
143 Gemini 2.5 Flash-Lite Preview (09-2025, no reasoning)Google
1,360
1,352–1,368
–
144 Gemini 1.5 Pro (002)Google
1,358
1,351–1,366
–
144 Qwen3.5-27BAlibaba
1,358
1,349–1,367
$0.27 / $2.16
144 Qwen3-VL-235B-A22B-InstructAlibaba
1,358
1,341–1,375
$0.30 / $1.50
147 MiniMax-M2.5MiniMax
1,357
1,349–1,365
$0.30 / $1.20
148 Qwen2.5-MaxAlibaba
1,353
1,345–1,362
–
148 Hy3 PreviewTencent
1,353
1,334–1,372
$0.18 / $0.60
150 MiMo-V2-Flash (no reasoning)Xiaomi · best of 2 settings
1,352
1,345–1,360
–
150 Trinity Large PreviewArcee AI
1,352
1,343–1,361
–
150 Nova Experimental Chat (2026-01-10)Amazon
1,352
1,327–1,377
–
153 GLM-4.6VZ.ai
1,351
1,322–1,379
$0.30 / $0.90
154 GPT-4.1 miniOpenAI
1,349
1,340–1,358
$0.40 / $1.60
154 DeepSeek-V3DeepSeek
1,349
1,339–1,359
$0.26 / $1.03
154 Gemma 3 27BGoogle
1,349
1,341–1,356
$0.12 / $0.20
157 Step 3.5 FlashStepFun
1,345
1,339–1,352
$0.10 / $0.30
157 Nova Experimental Chat (2026-02-10)Amazon
1,345
1,319–1,371
–
157 MiniMax-M2.1 PreviewMiniMax
1,345
1,333–1,356
–
157 Gemini Advanced (0514)Google
1,345
1,335–1,355
–
157 Gemini 2.0 Flash-Lite Preview (02-05)Google
1,345
1,335–1,354
–
157 Gemini 2.0 Flash (001)Google
1,345
1,337–1,352
–
163 Qwen3.5-35B-A3BAlibaba
1,342
1,333–1,351
$0.16 / $1.30
164 Nova Experimental Chat (12-10)Amazon
1,341
1,317–1,365
–
165 Qwen3-VL-235B-A22B-ThinkingAlibaba
1,340
1,322–1,358
$0.40 / $4
166 GPT-4o (2024-05-13)OpenAI
1,338
1,331–1,345
$5 / $15
167 Qwen3.5-FlashAlibaba
1,337
1,330–1,344
$0.065 / $0.26
167 o4-miniOpenAI
1,337
1,329–1,345
$1.10 / $4.40
167 Grok 3 Mini BetaxAI
1,337
1,325–1,348
–
170 GPT-5.4 nano (high reasoning)OpenAI
1,336
1,329–1,344
$0.20 / $1.25
170 Command ACohere
1,336
1,329–1,343
$2.50 / $10
170 Step-2 16K Exp (2024-12)StepFun
1,336
1,315–1,357
–
173 Gemma 3 12BGoogle
1,333
1,309–1,356
$0.05 / $0.15
174 Trinity Large ThinkingArcee AI
1,332
1,323–1,341
$0.25 / $0.80
174 Llama 3.1 Nemotron Ultra 253B v1NVIDIA
1,332
1,305–1,359
–
176 Grok 3 Mini (high reasoning)xAI
1,331
1,317–1,345
–
177 LongCat-Flash-ChatMeituan
1,330
1,314–1,346
–
178 GPT-4o (2024-08-06)OpenAI
1,327
1,319–1,336
$2.50 / $10
178 GLM-4.5-AirZ.ai
1,327
1,317–1,337
$0.14 / $0.86
180 Qwen3-Next-80B-A3B-ThinkingAlibaba
1,326
1,312–1,341
$0.15 / $1.20
181 GPT-5 mini (high reasoning)OpenAI
1,324
1,313–1,334
$0.25 / $2
181 Mistral Small 3.2 24BMistral AI
1,324
1,311–1,337
$0.094 / $0.25
183 Gemini 1.5 Pro (001)Google
1,323
1,315–1,331
–
183 GPT-4 TurboOpenAI
1,323
1,315–1,331
$10 / $30
185 MiniMax-M1MiniMax
1,319
1,310–1,329
$0.40 / $2.20
185 Qwen3-30B-A3B-Instruct-2507Alibaba
1,319
1,308–1,330
$0.09 / $0.30
187 Grok 2 (2024-08-13)xAI
1,317
1,310–1,325
–
188 Qwen-PlusAlibaba
1,316
1,298–1,335
$0.26 / $0.78
188 GLM-4-Plus (0111)Z.ai
1,316
1,297–1,335
–
190 Qwen3-Next-80B-A3B-InstructAlibaba
1,313
1,302–1,324
$0.10 / $1.10
191 o3-mini (high reasoning)OpenAI · best of 2 settings
1,312
1,301–1,323
$1.10 / $4.40
192 Inkling SmallThinking Machines
1,311
1,302–1,320
$0.45 / $1.20
193 Step 3StepFun
1,310
1,289–1,331
–
193 INTELLECT-3Prime Intellect
1,310
1,289–1,331
–
193 Solar Pro 4Upstage
1,310
1,298–1,321
$0.09 / $0.36
193 DeepSeek-V2.5-1210DeepSeek
1,310
1,293–1,327
–
193 GLM-4.5VZ.ai
1,310
1,286–1,333
$0.60 / $1.80
198 GPT-4.1 nanoOpenAI
1,307
1,289–1,326
$0.10 / $0.40
198 Llama 4 MaverickMeta
1,307
1,298–1,316
$0.27 / $0.85
198 GLM-4.7-FlashZ.ai
1,307
1,293–1,320
$0.06 / $0.40
201 Claude 3.5 HaikuAnthropic
1,306
1,299–1,312
–
202 Llama 3.1 405B Instruct (FP8)Meta
1,305
1,297–1,312
–
202 Claude 3.5 Sonnet (2024-06-20)Anthropic
1,305
1,297–1,312
–
202 Llama 3.3 Nemotron Super 49B v1.5NVIDIA
1,305
1,275–1,334
–
205 Qwen3-32BAlibaba
1,304
1,281–1,327
$0.14 / $0.40
206 Nemotron 3 Super 120B A12BNVIDIA
1,303
1,286–1,321
$0.085 / $0.40
206 Hunyuan Large (2025-02-10)Tencent
1,303
1,279–1,327
–
206 Hunyuan TurboS (2025-02-26)Tencent
1,303
1,277–1,329
–
209 Hunyuan Turbo (0110)Tencent
1,302
1,277–1,328
–
210 Llama 3.1 405B Instruct (BF16)Meta
1,301
1,294–1,309
–
211 Nova Experimental Chat (11-10)Amazon
1,300
1,290–1,309
–
212 Gemma 3n E4BGoogle
1,299
1,287–1,310
–
213 Gemini 1.5 Flash (002)Google
1,298
1,289–1,307
–
213 Llama 3.3 Nemotron Super 49B v1NVIDIA
1,298
1,270–1,325
–
215 Yi-Lightning01.AI
1,297
1,286–1,307
–
216 GPT-4o mini (2024-07-18)OpenAI
1,294
1,287–1,301
$0.15 / $0.60
216 Step-1o Turbo (2025-06)StepFun
1,294
1,274–1,313
–
218 QwQ-32BAlibaba
1,293
1,283–1,303
–
218 Olmo 3.1 32B InstructAi2
1,293
1,278–1,308
–
218 Gemma 2 27BGoogle
1,293
1,286–1,300
$0.65 / $0.65
221 Mercury 2Inception
1,292
1,266–1,319
$0.25 / $0.75
222 GLM-4-PlusZ.ai
1,290
1,279–1,300
–
223 Llama 4 ScoutMeta
1,289
1,279–1,299
$0.18 / $0.59
224 Claude 3 OpusAnthropic
1,288
1,281–1,295
–
224 Nova Experimental Chat (10-09)Amazon
1,288
1,256–1,320
–
226 Mistral Large 2 (2407)Mistral AI
1,287
1,279–1,296
$2 / $6
227 GPT-4 Turbo Preview (1106)OpenAI
1,286
1,278–1,294
–
228 Llama 3.3 70B InstructMeta
1,285
1,278–1,292
$0.59 / $0.79
228 MiniMax-M2MiniMax
1,285
1,265–1,306
$0.30 / $1.20
228 Qwen-Max (0919)Alibaba
1,285
1,273–1,297
–
231 Gemma 2 9B SimPOPrinceton NLP
1,284
1,269–1,299
–
231 Qwen3-30B-A3BAlibaba
1,284
1,273–1,294
$0.12 / $0.50
233 o1-miniOpenAI
1,281
1,273–1,289
–
233 Hunyuan Large VisionTencent
1,281
1,257–1,304
–
235 GPT-4 Turbo Preview (0125)OpenAI
1,280
1,272–1,289
–
235 Magistral Medium 1.0Mistral AI
1,280
1,263–1,297
–
237 Nova Experimental Chat (10-20)Amazon
1,279
1,263–1,294
–
238 gpt-oss-120bOpenAI
1,278
1,268–1,288
$0.15 / $0.60
239 Hunyuan Standard (2025-02-10)Tencent
1,277
1,253–1,301
–
239 Nemotron 3.5 Lightning 30B A3B (NVFP4)NVIDIA
1,277
1,264–1,289
–
241 Llama 3.1 Nemotron 70B InstructNVIDIA
1,276
1,258–1,295
–
241 Mistral Large 2.1 (2411)Mistral AI
1,276
1,267–1,285
–
243 Gemma 3 4BGoogle
1,275
1,252–1,297
$0.05 / $0.10
244 Qwen2.5-Plus (1127)Alibaba
1,273
1,259–1,288
–
245 Grok 2 Mini (2024-08-13)xAI
1,272
1,264–1,280
–
246 Athene 70B (0725)Nexusflow
1,271
1,259–1,282
–
246 Mistral Small 3.1 24BMistral AI
1,271
1,261–1,280
$0.35 / $0.56
248 Nova 2 LiteAmazon
1,269
1,254–1,284
$0.30 / $2.50
249 Reka Core (2024-09-04)Reka AI
1,268
1,249–1,286
–
250 GPT-4 (0613)OpenAI
1,267
1,258–1,275
–
251 Ling-flash-2.0Ant Group
1,266
1,246–1,287
–
252 Granite 4.2 30BIBM
1,265
1,240–1,291
–
252 DeepSeek-V2.5DeepSeek
1,265
1,254–1,276
–
254 Olmo 3 32B ThinkAi2
1,264
1,244–1,285
–
254 Gemini 1.5 Flash (001)Google
1,264
1,255–1,272
–
256 Command R+ (08-2024)Cohere
1,263
1,248–1,278
$2.50 / $10
257 Llama 3.1 Nemotron 51B InstructNVIDIA
1,262
1,236–1,287
–
257 Granite 4.1 8BIBM
1,262
1,237–1,287
–
259 Jamba 1.5 LargeAI21 Labs
1,261
1,245–1,278
–
260 Ring-flash-2.0Ant Group
1,259
1,239–1,279
–
260 Gemma 2 9BGoogle
1,259
1,251–1,266
–
262 Llama 3.1 70B InstructMeta
1,257
1,249–1,265
$0.40 / $0.40
263 GPT-4 (0314)OpenAI
1,255
1,245–1,265
–
264 Llama 3 70B InstructMeta
1,254
1,247–1,262
–
264 Qwen2.5-72B-InstructAlibaba
1,254
1,246–1,263
$0.36 / $0.40
266 Athene-V2-ChatNexusflow
1,253
1,243–1,263
–
267 Olmo 3.1 32B ThinkAi2
1,247
1,230–1,265
–
267 GPT-5 nano (high reasoning)OpenAI
1,247
1,227–1,267
$0.05 / $0.40
269 Llama 3.1 Tülu 3 70BAi2
1,246
1,222–1,271
–
269 Claude 3 SonnetAnthropic
1,246
1,238–1,254
–
271 Reka Flash (2024-09-04)Reka AI
1,243
1,225–1,260
–
272 Granite H SmallIBM
1,242
1,219–1,266
–
272 Nemotron 3 Nano 30B A3BNVIDIA
1,242
1,229–1,255
$0.05 / $0.20
274 GLM-4 (0520)Z.ai
1,241
1,226–1,256
–
275 gpt-oss-20bOpenAI
1,240
1,221–1,258
$0.03 / $0.15
276 Nemotron-4 340B InstructNVIDIA
1,238
1,226–1,250
–
276 Gemini 1.5 Flash-8B (001)Google
1,238
1,229–1,246
–
278 Nova Pro 1.0Amazon
1,237
1,227–1,246
$0.80 / $3.20
279 Command R+Cohere
1,236
1,228–1,245
–
280 Aya Expanse 32BCohere
1,228
1,218–1,238
–
281 Mistral Small 3Mistral AI
1,227
1,214–1,239
$0.05 / $0.08
282 OLMo 2 32B Instruct (0325)Ai2
1,224
1,199–1,248
–
283 Qwen2-72B-InstructAlibaba
1,223
1,213–1,232
–
284 Ministral 8B (2410)Mistral AI
1,222
1,200–1,244
–
285 Nova Lite 1.0Amazon
1,221
1,210–1,231
$0.06 / $0.24
286 MercuryInception
1,217
1,178–1,257
–
287 Claude 3 HaikuAnthropic
1,216
1,208–1,223
–
288 Jamba 1.5 MiniAI21 Labs
1,211
1,196–1,227
–
288 Mistral Large (2402)Mistral AI
1,211
1,202–1,221
–
290 Phi-4Microsoft
1,210
1,201–1,220
$0.07 / $0.14
291 Command R (08-2024)Cohere
1,209
1,194–1,225
$0.15 / $0.60
292 Qwen2.5-Coder-32B-InstructAlibaba
1,207
1,188–1,226
$0.66 / $1
293 DeepSeek-Coder-V2DeepSeek
1,205
1,192–1,219
–
293 Granite 4.2 3BIBM
1,205
1,177–1,233
–
295 Gemini Pro (Dev API)Google
1,204
1,190–1,218
–
296 WizardLM 70BMicrosoft
1,198
1,181–1,215
–
296 Nova Micro 1.0Amazon
1,198
1,187–1,208
$0.035 / $0.14
298 Granite 4.2 8BIBM
1,197
1,170–1,224
$0.06 / $0.25
298 Llama 3 8B InstructMeta
1,197
1,188–1,205
–
300 Llama 3.1 Tülu 3 8BAi2
1,196
1,171–1,221
–
301 Command RCohere
1,195
1,185–1,205
–
302 Hunyuan Standard 256KTencent
1,194
1,165–1,224
–
303 Qwen1.5-110B-ChatAlibaba
1,193
1,182–1,205
–
303 Mistral MediumMistral AI
1,193
1,182–1,204
–
305 GPT-3.5 Turbo (0125)OpenAI
1,191
1,182–1,200
–
305 Mixtral 8x22B InstructMistral AI
1,191
1,181–1,200
$2 / $6
307 Aya Expanse 8BCohere
1,189
1,174–1,204
–
308 Qwen1.5-72B-ChatAlibaba
1,188
1,178–1,198
–
309 Gemini ProGoogle
1,185
1,164–1,206
–
309 OpenChat 3.5OpenChat
1,185
1,167–1,203
–
311 Gemma 2 2BGoogle
1,183
1,175–1,192
–
312 Reka Flash 21B (2024-02-26, online)Reka AI
1,181
1,167–1,196
–
313 Llama 3.1 8B InstructMeta
1,177
1,169–1,186
$0.05 / $0.08
314 Reka Flash 21B (2024-02-26)Reka AI
1,175
1,163–1,187
–
315 Vicuna 33BLMSYS
1,172
1,160–1,184
–
316 Granite 3.1 8B InstructIBM
1,171
1,144–1,197
–
317 Zephyr ORPO 141B A35B v0.1Hugging Face
1,169
1,146–1,192
–
318 DBRX Instruct PreviewDatabricks
1,164
1,152–1,176
–
319 Mixtral 8x7B Instruct v0.1Mistral AI
1,159
1,151–1,168
–
319 Yi-1.5-34B-Chat01.AI
1,159
1,148–1,170
–
321 OpenHermes 2.5 Mistral 7BTeknium
1,158
1,136–1,180
–
322 GPT-3.5 Turbo (1106)OpenAI
1,156
1,141–1,171
–
323 Yi-34B-Chat01.AI
1,154
1,140–1,168
–
323 Falcon 180B ChatTII
1,154
1,117–1,190
–
325 Phi-3-medium-4k-instructMicrosoft
1,150
1,139–1,161
–
325 Gemma 1.1 7BGoogle
1,150
1,138–1,161
–
327 OpenChat 3.5 (0106)OpenChat
1,147
1,132–1,162
–
327 Granite 3.1 2B InstructIBM
1,147
1,118–1,176
–
329 Tülu 2 DPO 70BAi2
1,146
1,127–1,165
–
330 Llama 3.2 3B InstructMeta
1,145
1,125–1,164
$0.05 / $0.33
331 Snowflake Arctic InstructSnowflake
1,143
1,131–1,155
–
332 WizardLM 13BMicrosoft
1,141
1,123–1,158
–
333 Solar 10.7B Instruct v1.0Upstage
1,138
1,115–1,162
–
334 Qwen1.5-14B-ChatAlibaba
1,137
1,123–1,151
–
334 Granite 3.0 8B InstructIBM
1,137
1,115–1,158
–
336 Starling-LM-7B-alphaUC Berkeley
1,136
1,120–1,151
–
336 Zephyr 7B BetaHugging Face
1,136
1,120–1,152
–
338 Guanaco 33BTim Dettmers
1,135
1,108–1,163
–
338 DeepSeek LLM 67B ChatDeepSeek
1,135
1,114–1,157
–
338 Nous Hermes 2 Mixtral 8x7B DPONous Research
1,135
1,111–1,158
–
338 Dolphin 2.2.1 Mistral 7BCognitive Computations
1,135
1,101–1,168
–
342 Phi-3-small-8k-instructMicrosoft
1,133
1,120–1,147
–
342 MPT-30B-ChatMosaicML
1,133
1,104–1,161
–
344 Qwen1.5-32B-ChatAlibaba
1,131
1,119–1,144
–
345 InternLM2.5 20B ChatShanghai AI Lab
1,130
1,114–1,147
–
346 Vicuna 13BLMSYS
1,127
1,114–1,140
–
347 Llama 2 70B SteerLM ChatNVIDIA
1,119
1,094–1,145
–
348 Llama 2 70B ChatMeta
1,112
1,101–1,122
–
349 Starling-LM-7B-betaNexusflow
1,107
1,092–1,122
–
350 QwQ-32B-PreviewAlibaba
1,104
1,075–1,134
–
350 Mistral 7B Instruct v0.2Mistral AI
1,104
1,091–1,116
–
352 Qwen-14B-ChatAlibaba
1,103
1,082–1,123
–
353 Gemma 7BGoogle
1,101
1,084–1,119
–
354 Zephyr 7B AlphaHugging Face
1,100
1,071–1,130
–
354 Granite 3.0 2B InstructIBM
1,100
1,078–1,121
–
356 Phi-3-mini-4k-instruct (June 2024)Microsoft
1,096
1,081–1,111
–
357 Gemma 1.1 2BGoogle
1,094
1,077–1,111
–
357 Alpaca 13BStanford
1,094
1,072–1,116
–
359 Llama 2 13B ChatMeta
1,092
1,079–1,104
–
359 Mistral 7B InstructMistral AI
1,092
1,075–1,108
–
361 Vicuna 7BLMSYS
1,090
1,072–1,108
–
362 StripedHyena Nous 7BTogether AI
1,089
1,069–1,110
–
362 PaLM 2Google
1,089
1,071–1,106
–
364 Code Llama 34B InstructMeta
1,087
1,070–1,103
–
365 Phi-3-mini-128k-instructMicrosoft
1,086
1,071–1,100
–
366 Llama 3.2 1B InstructMeta
1,083
1,062–1,103
$0.027 / $0.20
367 Qwen1.5-7B-ChatAlibaba
1,081
1,058–1,104
–
368 SmolLM2 1.7B InstructHugging Face
1,079
1,045–1,114
–
368 Gemma 2BGoogle
1,079
1,056–1,102
–
370 Phi-3-mini-4k-instructMicrosoft
1,073
1,060–1,086
–
371 Llama 2 7B ChatMeta
1,072
1,059–1,086
–
372 GPT4All 13B SnoozyNomic AI
1,064
1,027–1,101
–
373 MPT-7B-ChatMosaicML
1,053
1,029–1,077
–
374 Qwen1.5-4B-ChatAlibaba
1,050
1,031–1,069
–
375 Koala 13BUC Berkeley
1,046
1,025–1,067
–
376 ChatGLM3-6BZ.ai
1,040
1,017–1,064
–
377 OpenAssistant Pythia 12BOpenAssistant
1,013
992–1,033
–
378 RWKV-4 Raven 14BRWKV
1,004
982–1,027
–
378 ChatGLM2-6BZ.ai
1,004
978–1,030
–
380 OLMo 7B InstructAi2
1,002
981–1,022
–
381 FastChat-T5 3BLMSYS
992
968–1,017
–
382 Dolly v2 12BDatabricks
970
943–998
–
383 ChatGLM-6BZ.ai
951
927–976
–
384 LLaMA 13BMeta
936
902–969
–
385 StableLM Tuned Alpha 7BStability AI
932
904–961
–

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. ≈ marks results whose 95% range overlaps the leader's: they cannot be told apart from it. Each model is shown at its best setting; show every setting. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Human preference on creative writing prompts.

What it does not measure

Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.