Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Human preference on creative writing prompts.

Last updated 2 Oct 2026

Results dated
2 Oct 2026
Results
411 configurations of 385 models
Unit
Arena rating
Licence
Creative Commons Attribution 4.0 International

Creative writing: Claude Fable 5

Top 15 of 411 results · Arena rating, higher is better · lines show the 95% range · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ Gemini 4 Argon (high reasoning)Google 1,519
  2. 2≈ Claude Opus 5.5 (high reasoning)Anthropic 1,516
  3. 3≈ Claude Fable 5 (high reasoning)Anthropic 1,503
  4. 4≈ Claude Opus 4.6 (high reasoning)Anthropic 1,501
  5. 5≈ Gemini 3.7 Flash (high reasoning)Google 1,493
  6. 6 Claude Opus 4.7 (high reasoning)Anthropic 1,489
  7. 7 Gemini 3.8 Flash (high reasoning)Google 1,485
  8. 8 Gemini 3 Pro PreviewGoogle 1,484
  9. 9 Claude Opus 4.7Anthropic 1,483
  10. 10 Claude Fable 5.1 (max reasoning)Anthropic 1,482
  11. 11 Gemini 3.1 Pro PreviewGoogle 1,480
  12. 12 Claude Opus 4.6Anthropic 1,479
  13. 13 Claude Opus 5 (high reasoning)Anthropic 1,471
  14. 13 Gemini 3.6 Flash (high reasoning)Google 1,471
  15. 13 Claude Opus 5 (max reasoning)Anthropic 1,471

Full results

Arena Text: creative writing, Arena rating, higher is better
#ModelCreative writing · 95% range
Arena rating, higher is better
Price
$ per million tokens, in / out
1≈ Gemini 4 Argon (high reasoning)Google
1,519
1,499–1,538
–
2≈ Claude Opus 5.5 (high reasoning)Anthropic
1,516
1,497–1,536
$4 / $20
3≈ Claude Fable 5 (high reasoning)Anthropic
1,503
1,495–1,510
$10 / $50
4≈ Claude Opus 4.6 (high reasoning)Anthropic
1,501
1,494–1,507
$5 / $25
5≈ Gemini 3.7 Flash (high reasoning)Google
1,493
1,484–1,503
$1.50 / $7.50
6 Claude Opus 4.7 (high reasoning)Anthropic
1,489
1,482–1,496
$5 / $25
7 Gemini 3.8 Flash (high reasoning)Google
1,485
1,476–1,494
$1.50 / $7.50
8 Gemini 3 Pro PreviewGoogle
1,484
1,476–1,493
–
9 Claude Opus 4.7Anthropic
1,483
1,476–1,490
$5 / $25
10 Claude Fable 5.1 (max reasoning)Anthropic
1,482
1,469–1,494
$10 / $50
11 Gemini 3.1 Pro PreviewGoogle
1,480
1,475–1,486
$2 / $12
12 Claude Opus 4.6Anthropic
1,479
1,473–1,485
$5 / $25
13 Claude Opus 5 (high reasoning)Anthropic
1,471
1,465–1,478
$5 / $25
13 Gemini 3.6 Flash (high reasoning)Google
1,471
1,463–1,479
$1.50 / $7.50
13 Claude Opus 5 (max reasoning)Anthropic
1,471
1,462–1,479
$5 / $25
16 Claude Opus 4.5 (high reasoning, 32k budget)Anthropic
1,470
1,461–1,478
$5 / $25
16 Claude Opus 4.8 (high reasoning)Anthropic
1,470
1,463–1,476
$5 / $25
18 Gemini 3.5 Flash (medium reasoning)Google
1,469
1,461–1,476
$1.50 / $9
19 GPT-5.6 Sol (extra-high reasoning)OpenAI
1,468
1,460–1,476
$4 / $20
19 Qwen3.8-Max (0902)Alibaba
1,468
1,458–1,478
$2 / $6
21 Muse SparkMeta
1,465
1,452–1,479
–
22 Gemini 3.5 Flash (high reasoning)Google
1,464
1,457–1,471
$1.50 / $9
23 Claude Opus 4.8Anthropic
1,463
1,457–1,470
$5 / $25
23 Grok 4.20 Beta 1xAI
1,463
1,453–1,473
–
25 Claude Opus 4.5Anthropic
1,462
1,456–1,468
$5 / $25
26 Kimi K3 (max reasoning)Moonshot AI
1,460
1,451–1,469
$3 / $15
26 GPT-6.1 Sol (max reasoning)OpenAI
1,460
1,436–1,483
$2 / $10
28 Muse Spark 1.3 (max reasoning)Meta
1,459
1,446–1,471
$1.25 / $4.25
29 Gemini 3 Flash PreviewGoogle
1,457
1,448–1,466
$0.50 / $3
30 GPT-5.5 InstantOpenAI
1,456
1,447–1,466
–
31 Muse Spark 1.2 (extra-high reasoning)Meta
1,455
1,435–1,476
$1.25 / $4.25
32 GLM-5.2 (max reasoning)Z.ai
1,454
1,447–1,462
$1.40 / $4.40
32 Claude Sonnet 4.5Anthropic
1,454
1,448–1,460
$3 / $15
32 GLM-5.3 (max reasoning)Z.ai
1,454
1,444–1,464
$1.40 / $4.40
35 GPT-5.5 (high reasoning)OpenAI
1,451
1,445–1,458
$5 / $30
36 Grok 4.5xAI
1,450
1,443–1,458
$2 / $6
36 Claude Sonnet 4.5 (high reasoning, 32k budget)Anthropic
1,450
1,444–1,456
$3 / $15
38 Claude Sonnet 4.6Anthropic
1,449
1,443–1,456
$3 / $15
38 GPT-6 Astra (max reasoning)OpenAI
1,449
1,435–1,463
$10 / $50
38 GLM-5Z.ai
1,449
1,440–1,458
$0.95 / $2.55
41 GPT-5.5OpenAI
1,448
1,442–1,455
$5 / $30
42 Claude Sonnet 5.5 (extra-high reasoning)Anthropic
1,447
1,424–1,471
$2 / $10
42 Muse Spark 1.1Meta
1,447
1,439–1,455
$1.25 / $4.25
42 Grok 4.20 Multi-Agent Beta (0309)xAI
1,447
1,440–1,454
–
42 Gemini 3 Flash Preview (minimal reasoning)Google
1,447
1,441–1,453
$0.50 / $3
46 GLM-5.1Z.ai
1,446
1,440–1,453
$1.38 / $4.40
46 MiMo-V2.6-ProXiaomi
1,446
1,425–1,467
$0.43 / $0.87
46 DeepSeek-V4-Pro (0423)DeepSeek
1,446
1,439–1,453
$1.42 / $2.83
46 Claude Opus 4.1 (16k reasoning budget)Anthropic
1,446
1,438–1,453
$15 / $75
50 Grok 4.20 Beta (0309, reasoning)xAI
1,445
1,439–1,452
–
50 Qwen3.5-Max-PreviewAlibaba
1,445
1,434–1,456
–
50 Grok 4.6 (high reasoning)xAI
1,445
1,436–1,454
$2 / $6
50 DeepSeek-V4-Pro (0813, high reasoning)DeepSeek
1,445
1,431–1,458
$1.32 / $3.96
54 Gemini 2.5 ProGoogle
1,443
1,438–1,448
$1.25 / $10
55 GPT-5.4 (high reasoning)OpenAI
1,442
1,436–1,449
$2.50 / $15
55 Claude Opus 4.1Anthropic
1,442
1,436–1,448
$15 / $75
55 GPT-6 Sol (max reasoning)OpenAI
1,442
1,427–1,456
$2 / $10
58 Qwen3.7-Max-PreviewAlibaba
1,441
1,414–1,469
–
58 DeepSeek-V4-Pro (0423, high reasoning)DeepSeek
1,441
1,433–1,448
$1.42 / $2.83
60 DeepSeek-V4.1-Flash (max reasoning)DeepSeek
1,440
1,425–1,455
$0.30 / $1.20
61 Qwen3.6-Max-PreviewAlibaba
1,437
1,414–1,459
$1.03 / $6.16
62 GPT-4.5 PreviewOpenAI
1,436
1,425–1,448
–
62 Gemini 3.5 Flash-LiteGoogle
1,436
1,428–1,444
$0.30 / $2.50
62 Grok 4.7 (extra-high reasoning)xAI
1,436
1,419–1,452
$2 / $6
65 Claude Sonnet 5 (high reasoning)Anthropic
1,435
1,428–1,442
$2 / $10
65 GPT-5.4OpenAI
1,435
1,428–1,441
$2.50 / $15
67 MiMo-V2.5-ProXiaomi
1,434
1,428–1,440
$0.43 / $0.87
67 GPT-5.2 Chat (2026-02-10)OpenAI
1,434
1,425–1,442
–
69 GLM-5.3-FlashZ.ai
1,433
1,424–1,443
$0.15 / $0.50
69 Kimi K2.6Moonshot AI
1,433
1,425–1,441
$0.95 / $4
71 Claude Opus 4 (16k reasoning budget)Anthropic
1,432
1,423–1,442
–
71 Qwen3.7-PlusAlibaba
1,432
1,424–1,439
$0.32 / $1.28
73 Grok 4.1xAI
1,431
1,425–1,438
–
74 GPT-5.1 (high reasoning)OpenAI
1,428
1,420–1,437
$1.25 / $10
74 ERNIE 5.1Baidu
1,428
1,419–1,436
–
76 Grok 4.1 ThinkingxAI
1,426
1,420–1,433
–
76 GPT-5.6 Terra (extra-high reasoning)OpenAI
1,426
1,418–1,433
$2 / $12
78 MiMo-V2-ProXiaomi
1,424
1,413–1,434
–
79 Kimi K2.5 (reasoning on)Moonshot AI
1,423
1,417–1,430
$0.57 / $2.85
80 Grok 4.3xAI
1,422
1,416–1,429
$1.25 / $2.50
80 Hy3Tencent
1,422
1,408–1,435
$0.14 / $0.58
80 ChatGPT-4o (2025-03-26)OpenAI
1,422
1,416–1,428
–
80 ERNIE 5.0 (0110)Baidu
1,422
1,413–1,430
–
84 Step 5 PreviewStepFun
1,421
1,396–1,445
–
85 ERNIE 5.0 Preview (1203)Baidu
1,420
1,405–1,435
–
85 Gemma 4 31BGoogle
1,420
1,401–1,439
$0.14 / $0.40
87 Claude Opus 4Anthropic
1,417
1,408–1,425
–
88 Grok 4.1 Fast (reasoning)xAI
1,414
1,408–1,421
–
89 Gemini 3.1 Flash-Lite PreviewGoogle
1,413
1,406–1,420
$0.25 / $1.50
89 GPT-6 Luna (max reasoning)OpenAI
1,413
1,400–1,427
$0.10 / $0.50
91 DeepSeek-V4-Flash (0423, high reasoning)DeepSeek
1,410
1,403–1,418
$0.14 / $0.28
91 DeepSeek-V3.2-ExpDeepSeek
1,410
1,395–1,424
$0.27 / $0.41
93 GPT-5.6 Luna (extra-high reasoning)OpenAI
1,409
1,401–1,417
$0.20 / $1.20
94 GPT-5.1OpenAI
1,408
1,400–1,416
$1.25 / $10
94 DeepSeek-V4-Flash (0423)DeepSeek
1,408
1,401–1,416
$0.14 / $0.28
94 GPT-5.3 ChatOpenAI
1,408
1,399–1,417
–
97 Qwen3.6-PlusAlibaba
1,407
1,399–1,415
$0.33 / $1.95
97 ERNIE 5.0 Preview (1022)Baidu
1,407
1,385–1,429
–
99 MiniMax-M3MiniMax
1,406
1,399–1,413
$0.30 / $1.20
99 GLM-5V-TurboZ.ai
1,406
1,391–1,420
$1.20 / $4
101 Qwen3.5-397B-A17BAlibaba
1,405
1,400–1,411
$0.55 / $3.50
101 DeepSeek-V3.1-TerminusDeepSeek
1,405
1,378–1,433
$0.27 / $1
101 GLM-4.7Z.ai
1,405
1,391–1,418
$0.54 / $1.98
104 Dola-Seed-2.0-ProByteDance
1,403
1,397–1,410
–
104 GPT-4.1OpenAI
1,403
1,395–1,410
$2 / $8
106 DeepSeek-V3.1 (reasoning on)DeepSeek
1,402
1,387–1,417
$0.55 / $1.65
107 GPT-5.4 mini (high reasoning)OpenAI
1,401
1,394–1,409
$0.75 / $4.50
107 DeepSeek-V3.2DeepSeek
1,401
1,393–1,408
$0.30 / $0.96
109 GLM-4.6Z.ai
1,400
1,392–1,409
$0.50 / $2
109 Mistral Medium 3.5Mistral AI
1,400
1,386–1,414
$1.50 / $7.50
109 Gemma 4 26B A4BGoogle
1,400
1,381–1,418
$0.10 / $0.30
112 Grok 3 Preview (02-24)xAI
1,398
1,389–1,407
–
113 Claude Sonnet 4 (32k reasoning budget)Anthropic
1,396
1,386–1,405
$3 / $15
114 MiMo-V2-OmniXiaomi
1,395
1,384–1,406
–
114 Grok 4 (0709)xAI
1,395
1,386–1,403
–
116 MiMo-V2.5Xiaomi
1,394
1,387–1,402
$0.17 / $0.34
116 Gemini 2.5 FlashGoogle
1,394
1,389–1,399
$0.30 / $2.50
118 Qwen3-Max-PreviewAlibaba
1,393
1,383–1,404
–
118 Claude 3.7 Sonnet (32k reasoning budget)Anthropic
1,393
1,385–1,402
–
118 Qwen3-MaxAlibaba
1,393
1,376–1,410
$0.78 / $3.90
121 DeepSeek-V3.2 (reasoning on)DeepSeek
1,392
1,384–1,399
$0.30 / $0.96
122 GPT-5.2OpenAI
1,391
1,385–1,398
$1.75 / $14
122 DeepSeek-V3.2-Exp (reasoning on)DeepSeek
1,391
1,374–1,409
$0.27 / $0.41
122 Kimi K2 Thinking TurboMoonshot AI
1,391
1,384–1,398
–
122 Kimi K2.5 (no reasoning)Moonshot AI
1,391
1,375–1,407
$0.57 / $2.85
126 DeepSeek-V3-0324DeepSeek
1,390
1,382–1,398
$0.25 / $1
126 DeepSeek-R1-0528DeepSeek
1,390
1,377–1,402
$0.50 / $2.18
126 LongCat-Flash-Chat (2602, experimental)Meituan
1,390
1,380–1,399
–
129 GPT-5 ChatOpenAI
1,389
1,380–1,399
–
129 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek
1,389
1,361–1,417
$0.27 / $1
131 DeepSeek-V3.1DeepSeek
1,388
1,375–1,402
$0.55 / $1.65
131 Claude Sonnet 4Anthropic
1,388
1,379–1,397
$3 / $15
131 Claude Haiku 4.5Anthropic
1,388
1,383–1,393
$1 / $5
134 GPT-5.2 (high reasoning)OpenAI
1,387
1,379–1,394
$1.75 / $14
135 MiMo-V2.6-FlashXiaomi
1,385
1,367–1,404
$0.14 / $0.28
136 Grok 4 Fast ChatxAI
1,384
1,364–1,404
–
137 Grok 4 Fast (reasoning)xAI
1,383
1,371–1,395
–
137 o3OpenAI
1,383
1,376–1,390
$2 / $8
139 Nemotron 3 Ultra 550B A55B (NVFP4)NVIDIA
1,382
1,368–1,395
–
140 o1OpenAI
1,381
1,373–1,390
$15 / $60
140 Kimi K2 (0905)Moonshot AI
1,381
1,366–1,397
$0.60 / $2.50
140 Gemini 2.5 Flash Preview (09-2025)Google
1,381
1,372–1,390
–
143 InklingThinking Machines
1,379
1,371–1,387
$0.95 / $4.05
144 Qwen3-235B-A22B-Instruct-2507Alibaba
1,378
1,373–1,384
$0.15 / $0.75
144 Mistral Medium 3.1Mistral AI
1,378
1,373–1,384
$0.40 / $2
144 Hunyuan Vision 1.5 ThinkingTencent
1,378
1,342–1,415
–
144 Claude 3.7 SonnetAnthropic
1,378
1,369–1,386
–
148 DeepSeek-R1DeepSeek
1,374
1,364–1,384
$0.70 / $2.50
148 Mistral Large 3Mistral AI
1,374
1,368–1,379
$0.50 / $1.50
150 Hunyuan T1 (2025-07-11)Tencent
1,373
1,350–1,397
–
150 GPT-5 (high reasoning)OpenAI
1,373
1,363–1,383
$1.25 / $10
152 Gemini 2.5 Flash-Lite Preview (06-17, reasoning on)Google
1,372
1,363–1,381
–
152 GLM-4.5Z.ai
1,372
1,361–1,383
$0.60 / $2.20
154 Kimi K2 (0711)Moonshot AI
1,371
1,361–1,381
$0.57 / $2.30
154 Claude 3.5 Sonnet (2024-10-22)Anthropic
1,371
1,365–1,377
–
154 Qwen3-235B-A22B-Thinking-2507Alibaba
1,371
1,352–1,389
$0.30 / $3
157 Qwen3-235B-A22B (no reasoning)Alibaba
1,366
1,357–1,375
$0.46 / $1.82
157 o1-previewOpenAI
1,366
1,356–1,376
–
157 Qwen3.5-122B-A10BAlibaba
1,366
1,356–1,375
$0.26 / $2.08
157 Muse GlimmerMeta
1,366
1,343–1,388
–
161 Mistral Medium 3Mistral AI
1,364
1,355–1,374
$0.40 / $2
161 MiniMax-M2.7MiniMax
1,364
1,358–1,370
$0.30 / $1.20
161 Qwen3.8-27BAlibaba
1,364
1,353–1,374
$0.50 / $3
164 Qwen3-Coder-480B-A35BAlibaba
1,363
1,352–1,373
$0.35 / $1.50
165 Hunyuan TurboS (2025-04-16)Tencent
1,361
1,346–1,377
–
166 Gemini 2.5 Flash-Lite Preview (09-2025, no reasoning)Google
1,360
1,352–1,368
–
167 Gemini 1.5 Pro (002)Google
1,358
1,351–1,366
–
167 Qwen3.5-27BAlibaba
1,358
1,349–1,367
$0.27 / $2.16
167 Qwen3-VL-235B-A22B-InstructAlibaba
1,358
1,341–1,375
$0.30 / $1.50
170 MiniMax-M2.5MiniMax
1,357
1,349–1,365
$0.30 / $1.20
171 Qwen2.5-MaxAlibaba
1,353
1,345–1,362
–
171 Hy3 PreviewTencent
1,353
1,334–1,372
$0.18 / $0.60
173 MiMo-V2-Flash (no reasoning)Xiaomi
1,352
1,345–1,360
–
173 Trinity Large PreviewArcee AI
1,352
1,343–1,361
–
173 Nova Experimental Chat (2026-01-10)Amazon
1,352
1,327–1,377
–
176 GLM-4.6VZ.ai
1,351
1,322–1,379
$0.30 / $0.90
177 GPT-4.1 miniOpenAI
1,349
1,340–1,358
$0.40 / $1.60
177 DeepSeek-V3DeepSeek
1,349
1,339–1,359
$0.26 / $1.03
177 Gemma 3 27BGoogle
1,349
1,341–1,356
$0.12 / $0.20
180 Step 3.5 FlashStepFun
1,345
1,339–1,352
$0.10 / $0.30
180 Nova Experimental Chat (2026-02-10)Amazon
1,345
1,319–1,371
–
180 MiniMax-M2.1 PreviewMiniMax
1,345
1,333–1,356
–
180 Gemini Advanced (0514)Google
1,345
1,335–1,355
–
180 Gemini 2.0 Flash-Lite Preview (02-05)Google
1,345
1,335–1,354
–
180 Gemini 2.0 Flash (001)Google
1,345
1,337–1,352
–
186 Qwen3.5-35B-A3BAlibaba
1,342
1,333–1,351
$0.16 / $1.30
187 Nova Experimental Chat (12-10)Amazon
1,341
1,317–1,365
–
188 Qwen3-VL-235B-A22B-ThinkingAlibaba
1,340
1,322–1,358
$0.40 / $4
189 GPT-4o (2024-05-13)OpenAI
1,338
1,331–1,345
$5 / $15
190 Qwen3.5-FlashAlibaba
1,337
1,330–1,344
$0.065 / $0.26
190 o4-miniOpenAI
1,337
1,329–1,345
$1.10 / $4.40
190 Grok 3 Mini BetaxAI
1,337
1,325–1,348
–
193 MiMo-V2-Flash (reasoning on)Xiaomi
1,336
1,322–1,350
–
193 GPT-5.4 nano (high reasoning)OpenAI
1,336
1,329–1,344
$0.20 / $1.25
193 Command ACohere
1,336
1,329–1,343
$2.50 / $10
193 Step-2 16K Exp (2024-12)StepFun
1,336
1,315–1,357
–
197 Gemma 3 12BGoogle
1,333
1,309–1,356
$0.05 / $0.15
198 Trinity Large ThinkingArcee AI
1,332
1,323–1,341
$0.25 / $0.80
198 Llama 3.1 Nemotron Ultra 253B v1NVIDIA
1,332
1,305–1,359
–
200 Grok 3 Mini (high reasoning)xAI
1,331
1,317–1,345
–
201 LongCat-Flash-ChatMeituan
1,330
1,314–1,346
–
202 GPT-4o (2024-08-06)OpenAI
1,327
1,319–1,336
$2.50 / $10
202 GLM-4.5-AirZ.ai
1,327
1,317–1,337
$0.14 / $0.86
204 Qwen3-Next-80B-A3B-ThinkingAlibaba
1,326
1,312–1,341
$0.15 / $1.20
205 GPT-5 mini (high reasoning)OpenAI
1,324
1,313–1,334
$0.25 / $2
205 Mistral Small 3.2 24BMistral AI
1,324
1,311–1,337
$0.094 / $0.25
207 Gemini 1.5 Pro (001)Google
1,323
1,315–1,331
–
207 Qwen3-235B-A22BAlibaba
1,323
1,313–1,333
$0.46 / $1.82
207 GPT-4 TurboOpenAI
1,323
1,315–1,331
$10 / $30
210 MiniMax-M1MiniMax
1,319
1,310–1,329
$0.40 / $2.20
210 Qwen3-30B-A3B-Instruct-2507Alibaba
1,319
1,308–1,330
$0.09 / $0.30
212 Grok 2 (2024-08-13)xAI
1,317
1,310–1,325
–
213 Qwen-PlusAlibaba
1,316
1,298–1,335
$0.26 / $0.78
213 GLM-4-Plus (0111)Z.ai
1,316
1,297–1,335
–
215 Qwen3-Next-80B-A3B-InstructAlibaba
1,313
1,302–1,324
$0.10 / $1.10
216 o3-mini (high reasoning)OpenAI
1,312
1,301–1,323
$1.10 / $4.40
217 Inkling SmallThinking Machines
1,311
1,302–1,320
$0.45 / $1.20
218 Step 3StepFun
1,310
1,289–1,331
–
218 INTELLECT-3Prime Intellect
1,310
1,289–1,331
–
218 Solar Pro 4Upstage
1,310
1,298–1,321
$0.09 / $0.36
218 DeepSeek-V2.5-1210DeepSeek
1,310
1,293–1,327
–
218 GLM-4.5VZ.ai
1,310
1,286–1,333
$0.60 / $1.80
223 GPT-4.1 nanoOpenAI
1,307
1,289–1,326
$0.10 / $0.40
223 Llama 4 MaverickMeta
1,307
1,298–1,316
$0.27 / $0.85
223 GLM-4.7-FlashZ.ai
1,307
1,293–1,320
$0.06 / $0.40
226 Claude 3.5 HaikuAnthropic
1,306
1,299–1,312
–
227 Llama 3.1 405B Instruct (FP8)Meta
1,305
1,297–1,312
–
227 Claude 3.5 Sonnet (2024-06-20)Anthropic
1,305
1,297–1,312
–
227 Llama 3.3 Nemotron Super 49B v1.5NVIDIA
1,305
1,275–1,334
–
230 Qwen3-32BAlibaba
1,304
1,281–1,327
$0.14 / $0.40
231 Nemotron 3 Super 120B A12BNVIDIA
1,303
1,286–1,321
$0.085 / $0.40
231 Hunyuan Large (2025-02-10)Tencent
1,303
1,279–1,327
–
231 Hunyuan TurboS (2025-02-26)Tencent
1,303
1,277–1,329
–
234 Hunyuan Turbo (0110)Tencent
1,302
1,277–1,328
–
235 Llama 3.1 405B Instruct (BF16)Meta
1,301
1,294–1,309
–
235 o3-miniOpenAI
1,301
1,294–1,308
$1.10 / $4.40
237 Nova Experimental Chat (11-10)Amazon
1,300
1,290–1,309
–
238 Gemma 3n E4BGoogle
1,299
1,287–1,310
–
239 Gemini 1.5 Flash (002)Google
1,298
1,289–1,307
–
239 Llama 3.3 Nemotron Super 49B v1NVIDIA
1,298
1,270–1,325
–
241 Yi-Lightning01.AI
1,297
1,286–1,307
–
242 GPT-4o mini (2024-07-18)OpenAI
1,294
1,287–1,301
$0.15 / $0.60
242 Step-1o Turbo (2025-06)StepFun
1,294
1,274–1,313
–
244 QwQ-32BAlibaba
1,293
1,283–1,303
–
244 Olmo 3.1 32B InstructAi2
1,293
1,278–1,308
–
244 Gemma 2 27BGoogle
1,293
1,286–1,300
$0.65 / $0.65
247 Mercury 2Inception
1,292
1,266–1,319
$0.25 / $0.75
248 GLM-4-PlusZ.ai
1,290
1,279–1,300
–
249 Llama 4 ScoutMeta
1,289
1,279–1,299
$0.18 / $0.59
250 Claude 3 OpusAnthropic
1,288
1,281–1,295
–
250 Nova Experimental Chat (10-09)Amazon
1,288
1,256–1,320
–
252 Mistral Large 2 (2407)Mistral AI
1,287
1,279–1,296
$2 / $6
253 GPT-4 Turbo Preview (1106)OpenAI
1,286
1,278–1,294
–
254 Llama 3.3 70B InstructMeta
1,285
1,278–1,292
$0.59 / $0.79
254 MiniMax-M2MiniMax
1,285
1,265–1,306
$0.30 / $1.20
254 Qwen-Max (0919)Alibaba
1,285
1,273–1,297
–
257 Gemma 2 9B SimPOPrinceton NLP
1,284
1,269–1,299
–
257 Qwen3-30B-A3BAlibaba
1,284
1,273–1,294
$0.12 / $0.50
259 o1-miniOpenAI
1,281
1,273–1,289
–
259 Hunyuan Large VisionTencent
1,281
1,257–1,304
–
261 GPT-4 Turbo Preview (0125)OpenAI
1,280
1,272–1,289
–
261 Magistral Medium 1.0Mistral AI
1,280
1,263–1,297
–
263 Nova Experimental Chat (10-20)Amazon
1,279
1,263–1,294
–
264 gpt-oss-120bOpenAI
1,278
1,268–1,288
$0.15 / $0.60
265 Hunyuan Standard (2025-02-10)Tencent
1,277
1,253–1,301
–
265 Nemotron 3.5 Lightning 30B A3B (NVFP4)NVIDIA
1,277
1,264–1,289
–
267 Llama 3.1 Nemotron 70B InstructNVIDIA
1,276
1,258–1,295
–
267 Mistral Large 2.1 (2411)Mistral AI
1,276
1,267–1,285
–
269 Gemma 3 4BGoogle
1,275
1,252–1,297
$0.05 / $0.10
270 Qwen2.5-Plus (1127)Alibaba
1,273
1,259–1,288
–
271 Grok 2 Mini (2024-08-13)xAI
1,272
1,264–1,280
–
272 Athene 70B (0725)Nexusflow
1,271
1,259–1,282
–
272 Mistral Small 3.1 24BMistral AI
1,271
1,261–1,280
$0.35 / $0.56
274 Nova 2 LiteAmazon
1,269
1,254–1,284
$0.30 / $2.50
275 Reka Core (2024-09-04)Reka AI
1,268
1,249–1,286
–
276 GPT-4 (0613)OpenAI
1,267
1,258–1,275
–
277 Ling-flash-2.0Ant Group
1,266
1,246–1,287
–
278 Granite 4.2 30BIBM
1,265
1,240–1,291
–
278 DeepSeek-V2.5DeepSeek
1,265
1,254–1,276
–
280 Olmo 3 32B ThinkAi2
1,264
1,244–1,285
–
280 Gemini 1.5 Flash (001)Google
1,264
1,255–1,272
–
282 Command R+ (08-2024)Cohere
1,263
1,248–1,278
$2.50 / $10
283 Llama 3.1 Nemotron 51B InstructNVIDIA
1,262
1,236–1,287
–
283 Granite 4.1 8BIBM
1,262
1,237–1,287
–
285 Jamba 1.5 LargeAI21 Labs
1,261
1,245–1,278
–
286 Ring-flash-2.0Ant Group
1,259
1,239–1,279
–
286 Gemma 2 9BGoogle
1,259
1,251–1,266
–
288 Llama 3.1 70B InstructMeta
1,257
1,249–1,265
$0.40 / $0.40
289 GPT-4 (0314)OpenAI
1,255
1,245–1,265
–
290 Llama 3 70B InstructMeta
1,254
1,247–1,262
–
290 Qwen2.5-72B-InstructAlibaba
1,254
1,246–1,263
$0.36 / $0.40
292 Athene-V2-ChatNexusflow
1,253
1,243–1,263
–
293 Olmo 3.1 32B ThinkAi2
1,247
1,230–1,265
–
293 GPT-5 nano (high reasoning)OpenAI
1,247
1,227–1,267
$0.05 / $0.40
295 Llama 3.1 Tülu 3 70BAi2
1,246
1,222–1,271
–
295 Claude 3 SonnetAnthropic
1,246
1,238–1,254
–
297 Reka Flash (2024-09-04)Reka AI
1,243
1,225–1,260
–
298 Granite H SmallIBM
1,242
1,219–1,266
–
298 Nemotron 3 Nano 30B A3BNVIDIA
1,242
1,229–1,255
$0.05 / $0.20
300 GLM-4 (0520)Z.ai
1,241
1,226–1,256
–
301 gpt-oss-20bOpenAI
1,240
1,221–1,258
$0.03 / $0.15
302 Nemotron-4 340B InstructNVIDIA
1,238
1,226–1,250
–
302 Gemini 1.5 Flash-8B (001)Google
1,238
1,229–1,246
–
304 Nova Pro 1.0Amazon
1,237
1,227–1,246
$0.80 / $3.20
305 Command R+Cohere
1,236
1,228–1,245
–
306 Aya Expanse 32BCohere
1,228
1,218–1,238
–
307 Mistral Small 3Mistral AI
1,227
1,214–1,239
$0.05 / $0.08
308 OLMo 2 32B Instruct (0325)Ai2
1,224
1,199–1,248
–
309 Qwen2-72B-InstructAlibaba
1,223
1,213–1,232
–
310 Ministral 8B (2410)Mistral AI
1,222
1,200–1,244
–
311 Nova Lite 1.0Amazon
1,221
1,210–1,231
$0.06 / $0.24
312 MercuryInception
1,217
1,178–1,257
–
313 Claude 3 HaikuAnthropic
1,216
1,208–1,223
–
314 Jamba 1.5 MiniAI21 Labs
1,211
1,196–1,227
–
314 Mistral Large (2402)Mistral AI
1,211
1,202–1,221
–
316 Phi-4Microsoft
1,210
1,201–1,220
$0.07 / $0.14
317 Command R (08-2024)Cohere
1,209
1,194–1,225
$0.15 / $0.60
318 Qwen2.5-Coder-32B-InstructAlibaba
1,207
1,188–1,226
$0.66 / $1
319 DeepSeek-Coder-V2DeepSeek
1,205
1,192–1,219
–
319 Granite 4.2 3BIBM
1,205
1,177–1,233
–
321 Gemini Pro (Dev API)Google
1,204
1,190–1,218
–
322 WizardLM 70BMicrosoft
1,198
1,181–1,215
–
322 Nova Micro 1.0Amazon
1,198
1,187–1,208
$0.035 / $0.14
324 Granite 4.2 8BIBM
1,197
1,170–1,224
$0.06 / $0.25
324 Llama 3 8B InstructMeta
1,197
1,188–1,205
–
326 Llama 3.1 Tülu 3 8BAi2
1,196
1,171–1,221
–
327 Command RCohere
1,195
1,185–1,205
–
328 Hunyuan Standard 256KTencent
1,194
1,165–1,224
–
329 Qwen1.5-110B-ChatAlibaba
1,193
1,182–1,205
–
329 Mistral MediumMistral AI
1,193
1,182–1,204
–
331 GPT-3.5 Turbo (0125)OpenAI
1,191
1,182–1,200
–
331 Mixtral 8x22B InstructMistral AI
1,191
1,181–1,200
$2 / $6
333 Aya Expanse 8BCohere
1,189
1,174–1,204
–
334 Qwen1.5-72B-ChatAlibaba
1,188
1,178–1,198
–
335 Gemini ProGoogle
1,185
1,164–1,206
–
335 OpenChat 3.5OpenChat
1,185
1,167–1,203
–
337 Gemma 2 2BGoogle
1,183
1,175–1,192
–
338 Reka Flash 21B (2024-02-26, online)Reka AI
1,181
1,167–1,196
–
339 Llama 3.1 8B InstructMeta
1,177
1,169–1,186
$0.05 / $0.08
340 Reka Flash 21B (2024-02-26)Reka AI
1,175
1,163–1,187
–
341 Vicuna 33BLMSYS
1,172
1,160–1,184
–
342 Granite 3.1 8B InstructIBM
1,171
1,144–1,197
–
343 Zephyr ORPO 141B A35B v0.1Hugging Face
1,169
1,146–1,192
–
344 DBRX Instruct PreviewDatabricks
1,164
1,152–1,176
–
345 Mixtral 8x7B Instruct v0.1Mistral AI
1,159
1,151–1,168
–
345 Yi-1.5-34B-Chat01.AI
1,159
1,148–1,170
–
347 OpenHermes 2.5 Mistral 7BTeknium
1,158
1,136–1,180
–
348 GPT-3.5 Turbo (1106)OpenAI
1,156
1,141–1,171
–
349 Yi-34B-Chat01.AI
1,154
1,140–1,168
–
349 Falcon 180B ChatTII
1,154
1,117–1,190
–
351 Phi-3-medium-4k-instructMicrosoft
1,150
1,139–1,161
–
351 Gemma 1.1 7BGoogle
1,150
1,138–1,161
–
353 OpenChat 3.5 (0106)OpenChat
1,147
1,132–1,162
–
353 Granite 3.1 2B InstructIBM
1,147
1,118–1,176
–
355 Tülu 2 DPO 70BAi2
1,146
1,127–1,165
–
356 Llama 3.2 3B InstructMeta
1,145
1,125–1,164
$0.05 / $0.33
357 Snowflake Arctic InstructSnowflake
1,143
1,131–1,155
–
358 WizardLM 13BMicrosoft
1,141
1,123–1,158
–
359 Solar 10.7B Instruct v1.0Upstage
1,138
1,115–1,162
–
360 Qwen1.5-14B-ChatAlibaba
1,137
1,123–1,151
–
360 Granite 3.0 8B InstructIBM
1,137
1,115–1,158
–
362 Starling-LM-7B-alphaUC Berkeley
1,136
1,120–1,151
–
362 Zephyr 7B BetaHugging Face
1,136
1,120–1,152
–
364 Guanaco 33BTim Dettmers
1,135
1,108–1,163
–
364 DeepSeek LLM 67B ChatDeepSeek
1,135
1,114–1,157
–
364 Nous Hermes 2 Mixtral 8x7B DPONous Research
1,135
1,111–1,158
–
364 Dolphin 2.2.1 Mistral 7BCognitive Computations
1,135
1,101–1,168
–
368 Phi-3-small-8k-instructMicrosoft
1,133
1,120–1,147
–
368 MPT-30B-ChatMosaicML
1,133
1,104–1,161
–
370 Qwen1.5-32B-ChatAlibaba
1,131
1,119–1,144
–
371 InternLM2.5 20B ChatShanghai AI Lab
1,130
1,114–1,147
–
372 Vicuna 13BLMSYS
1,127
1,114–1,140
–
373 Llama 2 70B SteerLM ChatNVIDIA
1,119
1,094–1,145
–
374 Llama 2 70B ChatMeta
1,112
1,101–1,122
–
375 Starling-LM-7B-betaNexusflow
1,107
1,092–1,122
–
376 QwQ-32B-PreviewAlibaba
1,104
1,075–1,134
–
376 Mistral 7B Instruct v0.2Mistral AI
1,104
1,091–1,116
–
378 Qwen-14B-ChatAlibaba
1,103
1,082–1,123
–
379 Gemma 7BGoogle
1,101
1,084–1,119
–
380 Zephyr 7B AlphaHugging Face
1,100
1,071–1,130
–
380 Granite 3.0 2B InstructIBM
1,100
1,078–1,121
–
382 Phi-3-mini-4k-instruct (June 2024)Microsoft
1,096
1,081–1,111
–
383 Gemma 1.1 2BGoogle
1,094
1,077–1,111
–
383 Alpaca 13BStanford
1,094
1,072–1,116
–
385 Llama 2 13B ChatMeta
1,092
1,079–1,104
–
385 Mistral 7B InstructMistral AI
1,092
1,075–1,108
–
387 Vicuna 7BLMSYS
1,090
1,072–1,108
–
388 StripedHyena Nous 7BTogether AI
1,089
1,069–1,110
–
388 PaLM 2Google
1,089
1,071–1,106
–
390 Code Llama 34B InstructMeta
1,087
1,070–1,103
–
391 Phi-3-mini-128k-instructMicrosoft
1,086
1,071–1,100
–
392 Llama 3.2 1B InstructMeta
1,083
1,062–1,103
$0.027 / $0.20
393 Qwen1.5-7B-ChatAlibaba
1,081
1,058–1,104
–
394 SmolLM2 1.7B InstructHugging Face
1,079
1,045–1,114
–
394 Gemma 2BGoogle
1,079
1,056–1,102
–
396 Phi-3-mini-4k-instructMicrosoft
1,073
1,060–1,086
–
397 Llama 2 7B ChatMeta
1,072
1,059–1,086
–
398 GPT4All 13B SnoozyNomic AI
1,064
1,027–1,101
–
399 MPT-7B-ChatMosaicML
1,053
1,029–1,077
–
400 Qwen1.5-4B-ChatAlibaba
1,050
1,031–1,069
–
401 Koala 13BUC Berkeley
1,046
1,025–1,067
–
402 ChatGLM3-6BZ.ai
1,040
1,017–1,064
–
403 OpenAssistant Pythia 12BOpenAssistant
1,013
992–1,033
–
404 RWKV-4 Raven 14BRWKV
1,004
982–1,027
–
404 ChatGLM2-6BZ.ai
1,004
978–1,030
–
406 OLMo 7B InstructAi2
1,002
981–1,022
–
407 FastChat-T5 3BLMSYS
992
968–1,017
–
408 Dolly v2 12BDatabricks
970
943–998
–
409 ChatGLM-6BZ.ai
951
927–976
–
410 LLaMA 13BMeta
936
902–969
–
411 StableLM Tuned Alpha 7BStability AI
932
904–961
–

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. ≈ marks results whose 95% range overlaps the leader's: they cannot be told apart from it. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Human preference on creative writing prompts.

What it does not measure

Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.