Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Human preference on prompts Arena classes as expert-level.

Last updated 2 Oct 2026

Results dated
2 Oct 2026
Results
363 configurations of 338 models
Unit
Arena rating
Licence
Creative Commons Attribution 4.0 International

Expert prompts: Claude Fable 5.1

Top 15 of 363 results · Arena rating, higher is better · lines show the 95% range · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ Claude Sonnet 5.5 (extra-high reasoning)Anthropic 1,556
  2. 2≈ Claude Fable 5 (high reasoning)Anthropic 1,549
  3. 3≈ Claude Opus 4.6 (high reasoning)Anthropic 1,547
  4. 4≈ GPT-6.1 Sol (max reasoning)OpenAI 1,543
  5. 4≈ Claude Opus 5 (high reasoning)Anthropic 1,543
  6. 6≈ Claude Opus 5.5 (high reasoning)Anthropic 1,541
  7. 7≈ Gemini 4 Argon (high reasoning)Google 1,539
  8. 8≈ GPT-5.6 Sol (extra-high reasoning)OpenAI 1,535
  9. 9≈ Claude Opus 4.7 (high reasoning)Anthropic 1,534
  10. 9≈ Claude Opus 4.7Anthropic 1,534
  11. 11≈ Kimi K3 (max reasoning)Moonshot AI 1,533
  12. 11≈ Claude Opus 4.6Anthropic 1,533
  13. 13≈ MiMo-V2.6-ProXiaomi 1,532
  14. 14≈ Claude Opus 5 (max reasoning)Anthropic 1,530
  15. 15≈ Muse Spark 1.3 (max reasoning)Meta 1,529
  16. 18≈ Claude Fable 5.1 (max reasoning)Anthropic 1,521

Full results

Arena Text: expert prompts, Arena rating, higher is better
#ModelExpert prompts · 95% range
Arena rating, higher is better
Price
$ per million tokens, in / out
1≈ Claude Sonnet 5.5 (extra-high reasoning)Anthropic
1,556
1,525–1,587
$2 / $10
2≈ Claude Fable 5 (high reasoning)Anthropic
1,549
1,540–1,559
$10 / $50
3≈ Claude Opus 4.6 (high reasoning)Anthropic
1,547
1,539–1,555
$5 / $25
4≈ GPT-6.1 Sol (max reasoning)OpenAI
1,543
1,511–1,576
$2 / $10
4≈ Claude Opus 5 (high reasoning)Anthropic
1,543
1,534–1,551
$5 / $25
6≈ Claude Opus 5.5 (high reasoning)Anthropic
1,541
1,514–1,567
$4 / $20
7≈ Gemini 4 Argon (high reasoning)Google
1,539
1,513–1,565
–
8≈ GPT-5.6 Sol (extra-high reasoning)OpenAI
1,535
1,525–1,544
$4 / $20
9≈ Claude Opus 4.7 (high reasoning)Anthropic
1,534
1,526–1,542
$5 / $25
9≈ Claude Opus 4.7Anthropic
1,534
1,526–1,542
$5 / $25
11≈ Kimi K3 (max reasoning)Moonshot AI
1,533
1,521–1,545
$3 / $15
11≈ Claude Opus 4.6Anthropic
1,533
1,526–1,541
$5 / $25
13≈ MiMo-V2.6-ProXiaomi
1,532
1,504–1,559
$0.43 / $0.87
14≈ Claude Opus 5 (max reasoning)Anthropic
1,530
1,518–1,541
$5 / $25
15≈ Muse Spark 1.3 (max reasoning)Meta
1,529
1,514–1,545
$1.25 / $4.25
16≈ Gemini 3.8 Flash (high reasoning)Google
1,526
1,515–1,537
$1.50 / $7.50
17≈ Claude Opus 4.8 (high reasoning)Anthropic
1,525
1,517–1,533
$5 / $25
18≈ Claude Fable 5.1 (max reasoning)Anthropic
1,521
1,502–1,540
$10 / $50
19≈ GLM-5.3 (max reasoning)Z.ai
1,518
1,504–1,532
$1.40 / $4.40
20≈ GPT-6 Astra (max reasoning)OpenAI
1,517
1,497–1,537
$10 / $50
20≈ DeepSeek-V4.1-Flash (max reasoning)DeepSeek
1,517
1,498–1,535
$0.30 / $1.20
22 Claude Opus 4.8Anthropic
1,516
1,508–1,524
$5 / $25
23 GPT-5.4 (high reasoning)OpenAI
1,515
1,507–1,524
$2.50 / $15
24 Qwen3.8-Max (0902)Alibaba
1,513
1,501–1,525
$2 / $6
25 GPT-5.5 (high reasoning)OpenAI
1,512
1,505–1,520
$5 / $30
25 Claude Sonnet 5 (high reasoning)Anthropic
1,512
1,503–1,521
$2 / $10
27 GPT-5.5OpenAI
1,511
1,504–1,519
$5 / $30
27 Gemini 3.7 Flash (high reasoning)Google
1,511
1,499–1,523
$1.50 / $7.50
27 GLM-5.3-FlashZ.ai
1,511
1,498–1,523
$0.15 / $0.50
30≈ MiMo-V2.6-FlashXiaomi
1,509
1,487–1,531
$0.14 / $0.28
30 Gemini 3.1 Pro PreviewGoogle
1,509
1,502–1,515
$2 / $12
32 MiMo-V2.5-ProXiaomi
1,507
1,499–1,515
$0.43 / $0.87
33 Claude Sonnet 4.6Anthropic
1,506
1,498–1,514
$3 / $15
33≈ Qwen3.7-Max-PreviewAlibaba
1,506
1,475–1,537
–
35 GPT-5.6 Terra (extra-high reasoning)OpenAI
1,504
1,495–1,514
$2 / $12
35 Muse Spark 1.1Meta
1,504
1,494–1,513
$1.25 / $4.25
37 Kimi K2.6Moonshot AI
1,503
1,493–1,513
$0.95 / $4
37 Claude Opus 4.5 (high reasoning, 32k budget)Anthropic
1,503
1,491–1,515
$5 / $25
37 Claude Opus 4.5Anthropic
1,503
1,494–1,511
$5 / $25
40≈ Muse Spark 1.2 (extra-high reasoning)Meta
1,502
1,476–1,528
$1.25 / $4.25
40 GLM-5.2 (max reasoning)Z.ai
1,502
1,493–1,511
$1.40 / $4.40
42 Qwen3.6-Max-PreviewAlibaba
1,500
1,475–1,525
$1.03 / $6.16
42 Gemini 3.6 Flash (high reasoning)Google
1,500
1,490–1,509
$1.50 / $7.50
42 Qwen3.5-Max-PreviewAlibaba
1,500
1,487–1,512
–
45 Gemini 3 Pro PreviewGoogle
1,498
1,487–1,509
–
45 GPT-6 Luna (max reasoning)OpenAI
1,498
1,480–1,516
$0.10 / $0.50
45 GPT-6 Sol (max reasoning)OpenAI
1,498
1,478–1,517
$2 / $10
45 Gemma 4 31BGoogle
1,498
1,473–1,523
$0.14 / $0.40
45 Claude Sonnet 4.5 (high reasoning, 32k budget)Anthropic
1,498
1,490–1,506
$3 / $15
50 Gemini 3 Flash PreviewGoogle
1,496
1,484–1,509
$0.50 / $3
51 ERNIE 5.1Baidu
1,495
1,484–1,505
–
52 GLM-5.1Z.ai
1,494
1,486–1,503
$1.38 / $4.40
52≈ Nova Experimental Chat (2026-02-10)Amazon
1,494
1,462–1,525
–
54 MiMo-V2-ProXiaomi
1,493
1,480–1,506
–
55 Gemini 3.5 Flash (high reasoning)Google
1,492
1,484–1,501
$1.50 / $9
55 GPT-5.4OpenAI
1,492
1,484–1,500
$2.50 / $15
55 Muse SparkMeta
1,492
1,476–1,508
–
58 GPT-5.6 Luna (extra-high reasoning)OpenAI
1,491
1,482–1,501
$0.20 / $1.20
58 Grok 4.6 (high reasoning)xAI
1,491
1,479–1,503
$2 / $6
58 GPT-5.2 Chat (2026-02-10)OpenAI
1,491
1,480–1,502
–
61 Grok 4.5xAI
1,489
1,479–1,498
$2 / $6
62 Hy3Tencent
1,488
1,471–1,506
$0.14 / $0.58
63 Gemini 3.5 Flash (medium reasoning)Google
1,487
1,479–1,496
$1.50 / $9
64 Qwen3.7-PlusAlibaba
1,485
1,476–1,494
$0.32 / $1.28
65 DeepSeek-V4-Pro (0813, high reasoning)DeepSeek
1,484
1,466–1,503
$1.32 / $3.96
66 Kimi K2.5 (reasoning on)Moonshot AI
1,483
1,475–1,490
$0.57 / $2.85
66 Claude Sonnet 4.5Anthropic
1,483
1,475–1,491
$3 / $15
68 Claude Opus 4.1 (16k reasoning budget)Anthropic
1,482
1,470–1,494
$15 / $75
68 GPT-5.1 (high reasoning)OpenAI
1,482
1,470–1,494
$1.25 / $10
68 GLM-5Z.ai
1,482
1,470–1,493
$0.95 / $2.55
71 Grok 4.20 Multi-Agent Beta (0309)xAI
1,481
1,473–1,489
–
72 GPT-5.4 mini (high reasoning)OpenAI
1,480
1,472–1,489
$0.75 / $4.50
72 DeepSeek-V4-Pro (0423)DeepSeek
1,480
1,471–1,488
$1.42 / $2.83
74 Grok 4.20 Beta (0309, reasoning)xAI
1,478
1,469–1,486
–
75 Dola-Seed-2.0-ProByteDance
1,477
1,470–1,485
–
75 Qwen3.5-397B-A17BAlibaba
1,477
1,470–1,484
$0.55 / $3.50
77 Qwen3.6-PlusAlibaba
1,476
1,467–1,485
$0.33 / $1.95
77 Grok 4.20 Beta 1xAI
1,476
1,463–1,488
–
77 DeepSeek-V4-Pro (0423, high reasoning)DeepSeek
1,476
1,467–1,484
$1.42 / $2.83
77 Qwen3.8-27BAlibaba
1,476
1,462–1,489
$0.50 / $3
81 Gemma 4 26B A4BGoogle
1,475
1,450–1,501
$0.10 / $0.30
81 GPT-5.2 (high reasoning)OpenAI
1,475
1,465–1,485
$1.75 / $14
83 GPT-5.5 InstantOpenAI
1,473
1,461–1,485
–
83 InklingThinking Machines
1,473
1,463–1,483
$0.95 / $4.05
85 Nova Experimental Chat (2026-01-10)Amazon
1,472
1,438–1,505
–
85 Step 5 PreviewStepFun
1,472
1,438–1,505
–
87 GPT-5.3 ChatOpenAI
1,471
1,459–1,482
–
88 Gemini 3.5 Flash-LiteGoogle
1,470
1,460–1,480
$0.30 / $2.50
88 MiMo-V2.5Xiaomi
1,470
1,461–1,479
$0.17 / $0.34
88 MiniMax-M3MiniMax
1,470
1,461–1,478
$0.30 / $1.20
91 LongCat-Flash-Chat (2602, experimental)Meituan
1,468
1,456–1,480
–
92 Nemotron 3 Ultra 550B A55B (NVFP4)NVIDIA
1,466
1,449–1,483
–
92 DeepSeek-V4-Flash (0423, high reasoning)DeepSeek
1,466
1,457–1,475
$0.14 / $0.28
92 Claude Opus 4.1Anthropic
1,466
1,457–1,475
$15 / $75
92 Grok 4.1 ThinkingxAI
1,466
1,457–1,475
–
92 Qwen3-235B-A22B-Thinking-2507Alibaba
1,466
1,437–1,494
$0.30 / $3
97 Qwen3-Max-PreviewAlibaba
1,465
1,448–1,481
–
98 Gemini 3 Flash Preview (minimal reasoning)Google
1,463
1,456–1,470
$0.50 / $3
98 GLM-5V-TurboZ.ai
1,463
1,446–1,480
$1.20 / $4
100 MiMo-V2-OmniXiaomi
1,462
1,448–1,475
–
101 Kimi K2 Thinking TurboMoonshot AI
1,460
1,451–1,469
–
102 GPT-5 (high reasoning)OpenAI
1,459
1,444–1,475
$1.25 / $10
102 GPT-5.2OpenAI
1,459
1,452–1,466
$1.75 / $14
104 Gemini 2.5 ProGoogle
1,458
1,452–1,465
$1.25 / $10
105 DeepSeek-V4-Flash (0423)DeepSeek
1,457
1,449–1,466
$0.14 / $0.28
105 Grok 4.7 (extra-high reasoning)xAI
1,457
1,434–1,480
$2 / $6
107 GPT-5.1OpenAI
1,455
1,444–1,466
$1.25 / $10
108 MiniMax-M2.7MiniMax
1,453
1,445–1,460
$0.30 / $1.20
109 DeepSeek-V3.2 (reasoning on)DeepSeek
1,452
1,441–1,464
$0.30 / $0.96
109 Grok 4.1xAI
1,452
1,444–1,461
–
109 Grok 4.3xAI
1,452
1,444–1,460
$1.25 / $2.50
112 Claude Haiku 4.5Anthropic
1,451
1,445–1,457
$1 / $5
112 Kimi K2.5 (no reasoning)Moonshot AI
1,451
1,428–1,474
$0.57 / $2.85
112 ERNIE 5.0 Preview (1203)Baidu
1,451
1,429–1,472
–
115 DeepSeek-V3.2-Exp (reasoning on)DeepSeek
1,449
1,420–1,477
$0.27 / $0.41
115 GLM-4.7Z.ai
1,449
1,428–1,469
$0.54 / $1.98
115 o3OpenAI
1,449
1,437–1,460
$2 / $8
118 DeepSeek-V3.2DeepSeek
1,448
1,438–1,458
$0.30 / $0.96
118 Claude Opus 4 (16k reasoning budget)Anthropic
1,448
1,434–1,462
–
120 Qwen3-235B-A22B-Instruct-2507Alibaba
1,447
1,440–1,455
$0.15 / $0.75
120 ERNIE 5.0 (0110)Baidu
1,447
1,436–1,458
–
122 Hy3 PreviewTencent
1,446
1,424–1,468
$0.18 / $0.60
123 Inkling SmallThinking Machines
1,445
1,434–1,456
$0.45 / $1.20
124 GLM-4.5Z.ai
1,444
1,426–1,461
$0.60 / $2.20
125 Gemini 3.1 Flash-Lite PreviewGoogle
1,443
1,435–1,452
$0.25 / $1.50
125 GPT-5 ChatOpenAI
1,443
1,428–1,458
–
127 Muse GlimmerMeta
1,442
1,414–1,470
–
127 Qwen3.5-122B-A10BAlibaba
1,442
1,430–1,453
$0.26 / $2.08
129 Qwen3.5-27BAlibaba
1,441
1,429–1,453
$0.27 / $2.16
130 Qwen3-VL-235B-A22B-InstructAlibaba
1,440
1,415–1,465
$0.30 / $1.50
131 Qwen3-VL-235B-A22B-ThinkingAlibaba
1,439
1,409–1,468
$0.40 / $4
131 GLM-4.6Z.ai
1,439
1,426–1,451
$0.50 / $2
133 Grok 4.1 Fast (reasoning)xAI
1,438
1,429–1,447
–
134 GPT-5.4 nano (high reasoning)OpenAI
1,437
1,429–1,446
$0.20 / $1.25
135 Solar Pro 4Upstage
1,436
1,422–1,451
$0.09 / $0.36
136 Mistral Medium 3.5Mistral AI
1,435
1,417–1,452
$1.50 / $7.50
137 Claude Opus 4Anthropic
1,434
1,421–1,447
–
138 Claude Sonnet 4 (32k reasoning budget)Anthropic
1,433
1,418–1,447
$3 / $15
139 ERNIE 5.0 Preview (1022)Baidu
1,432
1,399–1,466
–
139 DeepSeek-V3.1 (reasoning on)DeepSeek
1,432
1,407–1,457
$0.55 / $1.65
139 Gemini 2.5 Flash Preview (09-2025)Google
1,432
1,418–1,446
–
142 Grok 4 (0709)xAI
1,431
1,418–1,444
–
143 Step 3.5 FlashStepFun
1,429
1,420–1,437
$0.10 / $0.30
143 ChatGPT-4o (2025-03-26)OpenAI
1,429
1,420–1,438
–
143 GPT-4.5 PreviewOpenAI
1,429
1,405–1,452
–
146 LongCat-Flash-ChatMeituan
1,428
1,402–1,454
–
147 Mistral Large 3Mistral AI
1,427
1,420–1,434
$0.50 / $1.50
148 DeepSeek-V3.1DeepSeek
1,426
1,404–1,447
$0.55 / $1.65
149 MiniMax-M2.5MiniMax
1,425
1,415–1,435
$0.30 / $1.20
149 DeepSeek-V3.2-ExpDeepSeek
1,425
1,403–1,447
$0.27 / $0.41
151 Gemini 2.5 FlashGoogle
1,424
1,417–1,430
$0.30 / $2.50
152 Grok 4 Fast (reasoning)xAI
1,421
1,401–1,440
–
153 MiniMax-M2.1 PreviewMiniMax
1,420
1,403–1,437
–
154 Qwen3.5-35B-A3BAlibaba
1,419
1,408–1,431
$0.16 / $1.30
154 Qwen3.5-FlashAlibaba
1,419
1,411–1,428
$0.065 / $0.26
156 DeepSeek-R1-0528DeepSeek
1,418
1,399–1,438
$0.50 / $2.18
157 Qwen3-MaxAlibaba
1,417
1,391–1,444
$0.78 / $3.90
157 Kimi K2 (0905)Moonshot AI
1,417
1,393–1,441
$0.60 / $2.50
157 MiMo-V2-Flash (no reasoning)Xiaomi
1,417
1,407–1,427
–
160 Grok 4 Fast ChatxAI
1,416
1,382–1,451
–
160 Kimi K2 (0711)Moonshot AI
1,416
1,400–1,432
$0.57 / $2.30
162 MiMo-V2-Flash (reasoning on)Xiaomi
1,415
1,394–1,435
–
163 Claude 3.7 Sonnet (32k reasoning budget)Anthropic
1,412
1,398–1,426
–
164 Mistral Medium 3.1Mistral AI
1,410
1,402–1,418
$0.40 / $2
165 o4-miniOpenAI
1,408
1,395–1,420
$1.10 / $4.40
166 Hunyuan T1 (2025-07-11)Tencent
1,407
1,368–1,445
–
166 Grok 3 Mini (high reasoning)xAI
1,407
1,388–1,426
–
168 GPT-4.1OpenAI
1,405
1,393–1,417
$2 / $8
168 Nova Experimental Chat (11-10)Amazon
1,405
1,390–1,419
–
170 Trinity Large PreviewArcee AI
1,404
1,392–1,415
–
171 Grok 3 Preview (02-24)xAI
1,403
1,388–1,418
–
172 DeepSeek-R1DeepSeek
1,402
1,382–1,422
$0.70 / $2.50
172 o1OpenAI
1,402
1,385–1,419
$15 / $60
172 Claude Sonnet 4Anthropic
1,402
1,388–1,415
$3 / $15
175 GPT-5 mini (high reasoning)OpenAI
1,401
1,383–1,418
$0.25 / $2
176 Trinity Large ThinkingArcee AI
1,400
1,388–1,412
$0.25 / $0.80
176 Qwen3-Next-80B-A3B-InstructAlibaba
1,400
1,382–1,417
$0.10 / $1.10
178 GLM-4.5VZ.ai
1,399
1,357–1,441
$0.60 / $1.80
179 o3-mini (high reasoning)OpenAI
1,398
1,378–1,419
$1.10 / $4.40
179 DeepSeek-V3-0324DeepSeek
1,398
1,386–1,411
$0.25 / $1
179 Qwen3-32BAlibaba
1,398
1,360–1,435
$0.14 / $0.40
182 Nova Experimental Chat (12-10)Amazon
1,396
1,362–1,431
–
183 Qwen3-235B-A22B (no reasoning)Alibaba
1,395
1,382–1,409
$0.46 / $1.82
183 Granite 4.2 30BIBM
1,395
1,364–1,426
–
183 Qwen3-30B-A3B-Instruct-2507Alibaba
1,395
1,378–1,412
$0.09 / $0.30
186 GLM-4.6VZ.ai
1,394
1,351–1,437
$0.30 / $0.90
187 Qwen3-Next-80B-A3B-ThinkingAlibaba
1,393
1,370–1,416
$0.15 / $1.20
187 Nemotron 3 Super 120B A12BNVIDIA
1,393
1,370–1,415
$0.085 / $0.40
189 Claude 3.7 SonnetAnthropic
1,390
1,377–1,403
–
190 GLM-4.5-AirZ.ai
1,389
1,374–1,405
$0.14 / $0.86
191 Nemotron 3.5 Lightning 30B A3B (NVFP4)NVIDIA
1,388
1,373–1,404
–
192 GLM-4.7-FlashZ.ai
1,386
1,367–1,406
$0.06 / $0.40
192 Gemini 2.5 Flash-Lite Preview (09-2025, no reasoning)Google
1,386
1,375–1,397
–
194 Qwen3-235B-A22BAlibaba
1,385
1,369–1,401
$0.46 / $1.82
195 GPT-4.1 miniOpenAI
1,383
1,370–1,396
$0.40 / $1.60
195 o1-previewOpenAI
1,383
1,368–1,397
–
197 Grok 3 Mini BetaxAI
1,380
1,362–1,397
–
198 Qwen3-Coder-480B-A35BAlibaba
1,379
1,362–1,395
$0.35 / $1.50
198 Mistral Medium 3Mistral AI
1,379
1,365–1,392
$0.40 / $2
200 Claude 3.5 Sonnet (2024-10-22)Anthropic
1,372
1,362–1,381
–
201 Llama 3.3 Nemotron Super 49B v1.5NVIDIA
1,371
1,330–1,413
–
201 Gemini 2.5 Flash-Lite Preview (06-17, reasoning on)Google
1,371
1,356–1,386
–
201 Granite 4.2 8BIBM
1,371
1,337–1,405
$0.06 / $0.25
204 Qwen-PlusAlibaba
1,370
1,340–1,400
$0.26 / $0.78
205 Qwen2.5-MaxAlibaba
1,368
1,354–1,382
–
206 Ring-flash-2.0Ant Group
1,367
1,335–1,399
–
207 o3-miniOpenAI
1,366
1,354–1,377
$1.10 / $4.40
208 MiniMax-M1MiniMax
1,365
1,351–1,380
$0.40 / $2.20
208 Nova Experimental Chat (10-20)Amazon
1,365
1,343–1,386
–
210 Step 3StepFun
1,363
1,328–1,399
–
211 Ling-flash-2.0Ant Group
1,362
1,332–1,392
–
212 QwQ-32BAlibaba
1,361
1,344–1,377
–
213 gpt-oss-120bOpenAI
1,357
1,341–1,373
$0.15 / $0.60
214 Gemini 2.0 Flash (001)Google
1,356
1,343–1,369
–
215 INTELLECT-3Prime Intellect
1,354
1,323–1,386
–
215 GPT-5 nano (high reasoning)OpenAI
1,354
1,321–1,386
$0.05 / $0.40
217 Nova 2 LiteAmazon
1,352
1,332–1,373
$0.30 / $2.50
218 o1-miniOpenAI
1,351
1,340–1,363
–
219 MiniMax-M2MiniMax
1,350
1,317–1,382
$0.30 / $1.20
220 Hunyuan TurboS (2025-04-16)Tencent
1,349
1,325–1,373
–
221 Mercury 2Inception
1,348
1,312–1,384
$0.25 / $0.75
221 DeepSeek-V3DeepSeek
1,348
1,331–1,364
$0.26 / $1.03
223 Claude 3.5 Sonnet (2024-06-20)Anthropic
1,343
1,332–1,355
–
223 Command ACohere
1,343
1,332–1,355
$2.50 / $10
225 Qwen3-30B-A3BAlibaba
1,342
1,326–1,359
$0.12 / $0.50
226 Gemini 1.5 Pro (002)Google
1,339
1,328–1,350
–
226 Granite 4.1 8BIBM
1,339
1,307–1,370
–
228 Gemma 3 27BGoogle
1,337
1,324–1,350
$0.12 / $0.20
229 Olmo 3.1 32B ThinkAi2
1,335
1,310–1,361
–
230 Mistral Small 3.2 24BMistral AI
1,334
1,315–1,354
$0.094 / $0.25
231 Gemini 2.0 Flash-Lite Preview (02-05)Google
1,332
1,315–1,348
–
232 Yi-Lightning01.AI
1,331
1,316–1,345
–
232 Nemotron 3 Nano 30B A3BNVIDIA
1,331
1,312–1,350
$0.05 / $0.20
234 Granite 4.2 3BIBM
1,330
1,295–1,366
–
235 Olmo 3.1 32B InstructAi2
1,326
1,304–1,348
–
235 Llama 4 MaverickMeta
1,326
1,312–1,339
$0.27 / $0.85
235 Step-2 16K Exp (2024-12)StepFun
1,326
1,294–1,357
–
238 Llama 3.1 405B Instruct (FP8)Meta
1,325
1,313–1,336
–
239 Hunyuan Large (2025-02-10)Tencent
1,323
1,286–1,359
–
240 Qwen2.5-Plus (1127)Alibaba
1,321
1,300–1,343
–
241 Gemini 1.5 Pro (001)Google
1,320
1,308–1,332
–
242 Claude 3 OpusAnthropic
1,318
1,309–1,328
–
243 Claude 3.5 HaikuAnthropic
1,316
1,305–1,326
–
243 Grok 2 (2024-08-13)xAI
1,316
1,305–1,327
–
245 gpt-oss-20bOpenAI
1,314
1,287–1,342
$0.03 / $0.15
246 GPT-4.1 nanoOpenAI
1,313
1,283–1,344
$0.10 / $0.40
247 Granite H SmallIBM
1,312
1,276–1,348
–
247 Llama 3.1 405B Instruct (BF16)Meta
1,312
1,299–1,325
–
247 GLM-4-Plus (0111)Z.ai
1,312
1,282–1,341
–
250 GPT-4o (2024-08-06)OpenAI
1,311
1,299–1,324
$2.50 / $10
250 GPT-4o (2024-05-13)OpenAI
1,311
1,301–1,321
$5 / $15
250 Athene-V2-ChatNexusflow
1,311
1,296–1,326
–
253 Llama 4 ScoutMeta
1,308
1,293–1,324
$0.18 / $0.59
254 DeepSeek-V2.5-1210DeepSeek
1,306
1,280–1,332
–
255 GLM-4-PlusZ.ai
1,303
1,288–1,318
–
255 Olmo 3 32B ThinkAi2
1,303
1,267–1,339
–
257 Mistral Small 3.1 24BMistral AI
1,302
1,287–1,317
$0.35 / $0.56
257 Mistral Large 2 (2407)Mistral AI
1,302
1,289–1,314
$2 / $6
259 Athene 70B (0725)Nexusflow
1,301
1,282–1,320
–
260 Qwen-Max (0919)Alibaba
1,300
1,282–1,318
–
261 Gemini Advanced (0514)Google
1,298
1,284–1,313
–
261 Llama 3.3 70B InstructMeta
1,298
1,286–1,309
$0.59 / $0.79
263 Hunyuan Large VisionTencent
1,297
1,264–1,330
–
263 GPT-4 TurboOpenAI
1,297
1,286–1,308
$10 / $30
263 Step-1o Turbo (2025-06)StepFun
1,297
1,270–1,324
–
266 Llama 3.1 Nemotron 70B InstructNVIDIA
1,296
1,269–1,323
–
266 DeepSeek-V2.5DeepSeek
1,296
1,281–1,311
–
268 Grok 2 Mini (2024-08-13)xAI
1,295
1,283–1,307
–
268 Qwen2.5-72B-InstructAlibaba
1,295
1,283–1,307
$0.36 / $0.40
270 Magistral Medium 1.0Mistral AI
1,293
1,266–1,320
–
271 Reka Core (2024-09-04)Reka AI
1,291
1,266–1,316
–
272 GPT-4 Turbo Preview (1106)OpenAI
1,290
1,278–1,302
–
272 Hunyuan Standard (2025-02-10)Tencent
1,290
1,251–1,328
–
274 Jamba 1.5 LargeAI21 Labs
1,285
1,256–1,314
–
274 GPT-4o mini (2024-07-18)OpenAI
1,285
1,274–1,296
$0.15 / $0.60
274 Gemini 1.5 Flash (002)Google
1,285
1,272–1,298
–
277 GPT-4 Turbo Preview (0125)OpenAI
1,280
1,268–1,292
–
278 Gemma 3n E4BGoogle
1,279
1,261–1,298
–
279 Reka Flash (2024-09-04)Reka AI
1,277
1,254–1,300
–
280 Qwen2.5-Coder-32B-InstructAlibaba
1,276
1,244–1,309
$0.66 / $1
281 Llama 3.1 70B InstructMeta
1,275
1,264–1,287
$0.40 / $0.40
282 Gemma 3 12BGoogle
1,274
1,232–1,315
$0.05 / $0.15
283 Nova Pro 1.0Amazon
1,273
1,257–1,289
$0.80 / $3.20
283 Claude 3 SonnetAnthropic
1,273
1,261–1,285
–
285 Mistral Large 2.1 (2411)Mistral AI
1,271
1,256–1,286
–
286 Phi-4Microsoft
1,268
1,251–1,285
$0.07 / $0.14
286 Gemini 1.5 Flash (001)Google
1,268
1,256–1,281
–
288 Aya Expanse 32BCohere
1,267
1,253–1,281
–
289 DeepSeek-Coder-V2DeepSeek
1,266
1,245–1,288
–
290 GPT-4 (0314)OpenAI
1,264
1,249–1,280
–
291 Nova Lite 1.0Amazon
1,262
1,245–1,279
$0.06 / $0.24
292 Qwen2-72B-InstructAlibaba
1,261
1,246–1,275
–
293 Nemotron-4 340B InstructNVIDIA
1,260
1,242–1,279
–
293 Gemma 2 27BGoogle
1,260
1,250–1,271
$0.65 / $0.65
293 Gemma 3 4BGoogle
1,260
1,219–1,301
$0.05 / $0.10
296 Mistral Small 3Mistral AI
1,259
1,238–1,279
$0.05 / $0.08
297 GPT-4 (0613)OpenAI
1,253
1,240–1,266
–
297 Claude 3 HaikuAnthropic
1,253
1,242–1,264
–
297 Llama 3.1 Nemotron 51B InstructNVIDIA
1,253
1,220–1,286
–
300 GLM-4 (0520)Z.ai
1,252
1,227–1,277
–
301 Command R+ (08-2024)Cohere
1,250
1,225–1,274
$2.50 / $10
302 Llama 3 70B InstructMeta
1,245
1,234–1,256
–
303 Nova Micro 1.0Amazon
1,244
1,227–1,261
$0.035 / $0.14
304 Gemini 1.5 Flash-8B (001)Google
1,243
1,230–1,256
–
305 Command R+Cohere
1,242
1,229–1,254
–
306 Gemma 2 9BGoogle
1,238
1,226–1,249
–
307 Ministral 8B (2410)Mistral AI
1,237
1,207–1,268
–
308 Qwen1.5-72B-ChatAlibaba
1,236
1,220–1,251
–
309 Granite 3.1 8B InstructIBM
1,235
1,200–1,271
–
309 Gemma 2 9B SimPOPrinceton NLP
1,235
1,206–1,265
–
311 Aya Expanse 8BCohere
1,232
1,208–1,256
–
312 Command R (08-2024)Cohere
1,231
1,209–1,253
$0.15 / $0.60
313 Mistral Large (2402)Mistral AI
1,229
1,215–1,243
–
313 Qwen1.5-110B-ChatAlibaba
1,229
1,212–1,245
–
315 Mistral MediumMistral AI
1,226
1,209–1,243
–
315 Reka Flash 21B (2024-02-26, online)Reka AI
1,226
1,205–1,247
–
317 Qwen1.5-32B-ChatAlibaba
1,223
1,205–1,242
–
318 Yi-1.5-34B-Chat01.AI
1,221
1,202–1,239
–
319 InternLM2.5 20B ChatShanghai AI Lab
1,220
1,198–1,242
–
320 Mixtral 8x22B InstructMistral AI
1,218
1,204–1,232
$2 / $6
320 Granite 3.1 2B InstructIBM
1,218
1,183–1,254
–
322 Reka Flash 21B (2024-02-26)Reka AI
1,216
1,199–1,233
–
323 GPT-3.5 Turbo (1106)OpenAI
1,215
1,187–1,243
–
323 Command RCohere
1,215
1,201–1,228
–
325 Jamba 1.5 MiniAI21 Labs
1,213
1,182–1,244
–
326 Llama 3 8B InstructMeta
1,211
1,199–1,223
–
327 Phi-3-medium-4k-instructMicrosoft
1,205
1,187–1,223
–
327 Granite 3.0 8B InstructIBM
1,205
1,174–1,236
–
329 Llama 3.1 8B InstructMeta
1,198
1,185–1,210
$0.05 / $0.08
330 GPT-3.5 Turbo (0125)OpenAI
1,197
1,185–1,210
–
330 Mixtral 8x7B Instruct v0.1Mistral AI
1,197
1,184–1,210
–
332 Qwen1.5-14B-ChatAlibaba
1,194
1,174–1,214
–
333 DBRX Instruct PreviewDatabricks
1,191
1,174–1,207
–
334 Gemini Pro (Dev API)Google
1,180
1,156–1,203
–
335 Llama 3.2 3B InstructMeta
1,179
1,154–1,203
$0.05 / $0.33
336 Granite 3.0 2B InstructIBM
1,177
1,149–1,206
–
337 Gemma 1.1 7BGoogle
1,175
1,157–1,192
–
338 Starling-LM-7B-betaNexusflow
1,174
1,154–1,194
–
339 Zephyr ORPO 141B A35B v0.1Hugging Face
1,172
1,133–1,212
–
339 Gemma 2 2BGoogle
1,172
1,159–1,185
–
341 Phi-3-small-8k-instructMicrosoft
1,167
1,147–1,186
–
342 Snowflake Arctic InstructSnowflake
1,165
1,148–1,182
–
343 OpenChat 3.5OpenChat
1,164
1,123–1,205
–
343 OpenChat 3.5 (0106)OpenChat
1,164
1,141–1,187
–
345 Yi-34B-Chat01.AI
1,156
1,131–1,181
–
346 QwQ-32B-PreviewAlibaba
1,153
1,118–1,189
–
347 Phi-3-mini-4k-instruct (June 2024)Microsoft
1,149
1,123–1,175
–
348 Qwen1.5-7B-ChatAlibaba
1,148
1,112–1,185
–
349 Vicuna 33BLMSYS
1,146
1,120–1,173
–
350 Phi-3-mini-4k-instructMicrosoft
1,140
1,121–1,160
–
350 Mistral 7B Instruct v0.2Mistral AI
1,140
1,120–1,161
–
352 Starling-LM-7B-alphaUC Berkeley
1,138
1,107–1,168
–
353 Llama 2 70B ChatMeta
1,137
1,120–1,155
–
354 Llama 2 7B ChatMeta
1,134
1,106–1,162
–
355 Llama 2 13B ChatMeta
1,130
1,104–1,155
–
356 Gemma 7BGoogle
1,122
1,090–1,153
–
357 Gemma 1.1 2BGoogle
1,117
1,092–1,143
–
357 Vicuna 13BLMSYS
1,117
1,086–1,148
–
359 Zephyr 7B BetaHugging Face
1,113
1,075–1,151
–
360 Qwen1.5-4B-ChatAlibaba
1,108
1,077–1,139
–
361 Phi-3-mini-128k-instructMicrosoft
1,098
1,078–1,119
–
362 Mistral 7B InstructMistral AI
1,085
1,046–1,124
–
363 Llama 3.2 1B InstructMeta
1,082
1,054–1,110
$0.027 / $0.20

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. ≈ marks results whose 95% range overlaps the leader's: they cannot be told apart from it. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Human preference on prompts Arena classes as expert-level.

What it does not measure

Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.