Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Human preference on prompts that set explicit instructions.

Last updated 2 Oct 2026

Results dated
2 Oct 2026
Results
413 configurations of 387 models
Unit
Arena rating
Licence
Creative Commons Attribution 4.0 International

Instruction following: Gemini 4 Argon

Top 15 of 413 results · Arena rating, higher is better · lines show the 95% range · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ Gemini 4 Argon (high reasoning)Google 1,528
  2. 2≈ Claude Opus 5.5 (high reasoning)Anthropic 1,517
  3. 3≈ Claude Opus 4.6 (high reasoning)Anthropic 1,514
  4. 4≈ Claude Fable 5 (high reasoning)Anthropic 1,509
  5. 5 Claude Opus 4.7 (high reasoning)Anthropic 1,502
  6. 6 Claude Opus 4.6Anthropic 1,499
  7. 7 Claude Fable 5.1 (max reasoning)Anthropic 1,497
  8. 8 Claude Opus 5 (high reasoning)Anthropic 1,496
  9. 9 Claude Opus 4.7Anthropic 1,493
  10. 9 Claude Opus 5 (max reasoning)Anthropic 1,493
  11. 11 Claude Opus 4.8 (high reasoning)Anthropic 1,491
  12. 12 GPT-5.6 Sol (extra-high reasoning)OpenAI 1,490
  13. 12 GPT-6.1 Sol (max reasoning)OpenAI 1,490
  14. 14 Kimi K3 (max reasoning)Moonshot AI 1,488
  15. 15 Gemini 3.8 Flash (high reasoning)Google 1,486

Full results

Arena Text: instruction following, Arena rating, higher is better
#ModelInstruction following · 95% range
Arena rating, higher is better
Price
$ per million tokens, in / out
1≈ Gemini 4 Argon (high reasoning)Google
1,528
1,514–1,543
–
2≈ Claude Opus 5.5 (high reasoning)Anthropic
1,517
1,502–1,533
$4 / $20
3≈ Claude Opus 4.6 (high reasoning)Anthropic
1,514
1,508–1,519
$5 / $25
4≈ Claude Fable 5 (high reasoning)Anthropic
1,509
1,503–1,515
$10 / $50
5 Claude Opus 4.7 (high reasoning)Anthropic
1,502
1,497–1,508
$5 / $25
6 Claude Opus 4.6Anthropic
1,499
1,494–1,504
$5 / $25
7 Claude Fable 5.1 (max reasoning)Anthropic
1,497
1,487–1,508
$10 / $50
8 Claude Opus 5 (high reasoning)Anthropic
1,496
1,490–1,502
$5 / $25
9 Claude Opus 4.7Anthropic
1,493
1,488–1,498
$5 / $25
9 Claude Opus 5 (max reasoning)Anthropic
1,493
1,486–1,500
$5 / $25
11 Claude Opus 4.8 (high reasoning)Anthropic
1,491
1,485–1,496
$5 / $25
12 GPT-5.6 Sol (extra-high reasoning)OpenAI
1,490
1,484–1,497
$4 / $20
12 GPT-6.1 Sol (max reasoning)OpenAI
1,490
1,471–1,509
$2 / $10
14 Kimi K3 (max reasoning)Moonshot AI
1,488
1,481–1,495
$3 / $15
15 Gemini 3.8 Flash (high reasoning)Google
1,486
1,479–1,493
$1.50 / $7.50
15 Muse Spark 1.3 (max reasoning)Meta
1,486
1,476–1,495
$1.25 / $4.25
17 Gemini 3.7 Flash (high reasoning)Google
1,485
1,477–1,492
$1.50 / $7.50
18 Claude Sonnet 5.5 (extra-high reasoning)Anthropic
1,484
1,466–1,502
$2 / $10
19 Claude Opus 4.5 (high reasoning, 32k budget)Anthropic
1,483
1,476–1,490
$5 / $25
20 Gemini 3.1 Pro PreviewGoogle
1,480
1,475–1,484
$2 / $12
21 DeepSeek-V4.1-Flash (max reasoning)DeepSeek
1,479
1,469–1,490
$0.30 / $1.20
22 Claude Opus 4.8Anthropic
1,478
1,473–1,484
$5 / $25
22 GPT-5.5 (high reasoning)OpenAI
1,478
1,472–1,483
$5 / $30
24 Muse Spark 1.2 (extra-high reasoning)Meta
1,477
1,462–1,492
$1.25 / $4.25
25 GLM-5.3 (max reasoning)Z.ai
1,476
1,467–1,484
$1.40 / $4.40
26 GPT-5.5OpenAI
1,475
1,470–1,480
$5 / $30
26 Claude Opus 4.5Anthropic
1,475
1,470–1,480
$5 / $25
28 Claude Sonnet 4.6Anthropic
1,474
1,469–1,480
$3 / $15
28 MiMo-V2.6-ProXiaomi
1,474
1,458–1,490
$0.43 / $0.87
28 GPT-6 Astra (max reasoning)OpenAI
1,474
1,463–1,485
$10 / $50
28 Qwen3.8-Max (0902)Alibaba
1,474
1,466–1,481
$2 / $6
32 Gemini 3 Pro PreviewGoogle
1,473
1,466–1,479
–
33 Muse Spark 1.1Meta
1,472
1,466–1,479
$1.25 / $4.25
33 Gemini 3.6 Flash (high reasoning)Google
1,472
1,466–1,479
$1.50 / $7.50
33 GLM-5.3-FlashZ.ai
1,472
1,464–1,479
$0.15 / $0.50
36 GPT-5.4 (high reasoning)OpenAI
1,471
1,466–1,477
$2.50 / $15
37 GLM-5.2 (max reasoning)Z.ai
1,470
1,464–1,475
$1.40 / $4.40
38 MiMo-V2.5-ProXiaomi
1,469
1,464–1,474
$0.43 / $0.87
39 Claude Sonnet 5 (high reasoning)Anthropic
1,467
1,461–1,472
$2 / $10
39 Qwen3.5-Max-PreviewAlibaba
1,467
1,459–1,474
–
41 Gemini 3.5 Flash (medium reasoning)Google
1,466
1,460–1,471
$1.50 / $9
42 Qwen3.7-Max-PreviewAlibaba
1,465
1,449–1,482
–
42 Gemini 3.5 Flash (high reasoning)Google
1,465
1,459–1,471
$1.50 / $9
42 Muse SparkMeta
1,465
1,456–1,474
–
45 Claude Sonnet 4.5 (high reasoning, 32k budget)Anthropic
1,464
1,459–1,468
$3 / $15
46 Claude Sonnet 4.5Anthropic
1,462
1,457–1,466
$3 / $15
47 GPT-5.6 Terra (extra-high reasoning)OpenAI
1,461
1,455–1,468
$2 / $12
47 DeepSeek-V4-Pro (0813, high reasoning)DeepSeek
1,461
1,451–1,471
$1.32 / $3.96
47 Grok 4.5xAI
1,461
1,455–1,467
$2 / $6
50 GLM-5.1Z.ai
1,460
1,455–1,465
$1.38 / $4.40
50 GPT-5.4OpenAI
1,460
1,455–1,465
$2.50 / $15
52 Claude Opus 4.1 (16k reasoning budget)Anthropic
1,458
1,452–1,464
$15 / $75
53 Gemini 3 Flash PreviewGoogle
1,457
1,450–1,465
$0.50 / $3
53 GPT-5.2 Chat (2026-02-10)OpenAI
1,457
1,450–1,463
–
55 GPT-5.5 InstantOpenAI
1,456
1,449–1,463
–
55 Kimi K2.6Moonshot AI
1,456
1,449–1,462
$0.95 / $4
57 Claude Opus 4.1Anthropic
1,454
1,450–1,459
$15 / $75
58 MiMo-V2.6-FlashXiaomi
1,453
1,440–1,466
$0.14 / $0.28
58 Gemma 4 31BGoogle
1,453
1,439–1,467
$0.14 / $0.40
58 GPT-6 Sol (max reasoning)OpenAI
1,453
1,442–1,464
$2 / $10
58 ERNIE 5.1Baidu
1,453
1,446–1,459
–
62 DeepSeek-V4-Pro (0423)DeepSeek
1,452
1,447–1,458
$1.42 / $2.83
63 Grok 4.6 (high reasoning)xAI
1,451
1,443–1,458
$2 / $6
64 GPT-5.1 (high reasoning)OpenAI
1,450
1,444–1,457
$1.25 / $10
64 Grok 4.20 Beta 1xAI
1,450
1,443–1,457
–
64 Qwen3.6-Max-PreviewAlibaba
1,450
1,436–1,463
$1.03 / $6.16
67 DeepSeek-V4-Pro (0423, high reasoning)DeepSeek
1,448
1,442–1,454
$1.42 / $2.83
67 GPT-6 Luna (max reasoning)OpenAI
1,448
1,438–1,458
$0.10 / $0.50
69 Qwen3.7-PlusAlibaba
1,447
1,441–1,453
$0.32 / $1.28
69 Grok 4.20 Beta (0309, reasoning)xAI
1,447
1,441–1,452
–
71 GLM-5Z.ai
1,446
1,440–1,453
$0.95 / $2.55
71 Hy3Tencent
1,446
1,436–1,456
$0.14 / $0.58
73 Step 5 PreviewStepFun
1,445
1,427–1,464
–
73 GPT-5.6 Luna (extra-high reasoning)OpenAI
1,445
1,439–1,451
$0.20 / $1.20
73 MiMo-V2-ProXiaomi
1,445
1,437–1,452
–
76 Claude Opus 4 (16k reasoning budget)Anthropic
1,444
1,437–1,451
–
76 Grok 4.20 Multi-Agent Beta (0309)xAI
1,444
1,439–1,449
–
78 Gemini 3.5 Flash-LiteGoogle
1,443
1,437–1,449
$0.30 / $2.50
78 Gemini 3 Flash Preview (minimal reasoning)Google
1,443
1,438–1,448
$0.50 / $3
80 Grok 4.7 (extra-high reasoning)xAI
1,440
1,427–1,453
$2 / $6
81 Kimi K2.5 (reasoning on)Moonshot AI
1,439
1,434–1,444
$0.57 / $2.85
82 Gemini 2.5 ProGoogle
1,438
1,434–1,442
$1.25 / $10
83 Gemma 4 26B A4BGoogle
1,437
1,424–1,451
$0.10 / $0.30
83 GPT-4.5 PreviewOpenAI
1,437
1,428–1,445
–
85 GPT-5.3 ChatOpenAI
1,436
1,430–1,443
–
86 Qwen3.6-PlusAlibaba
1,435
1,430–1,441
$0.33 / $1.95
86 Dola-Seed-2.0-ProByteDance
1,435
1,430–1,440
–
88 MiniMax-M3MiniMax
1,434
1,429–1,440
$0.30 / $1.20
88 GPT-5.4 mini (high reasoning)OpenAI
1,434
1,428–1,439
$0.75 / $4.50
88 Qwen3.5-397B-A17BAlibaba
1,434
1,429–1,438
$0.55 / $3.50
91 DeepSeek-V4-Flash (0423, high reasoning)DeepSeek
1,433
1,428–1,439
$0.14 / $0.28
92 Kimi K2.5 (no reasoning)Moonshot AI
1,432
1,420–1,444
$0.57 / $2.85
93 Grok 4.1xAI
1,431
1,426–1,436
–
93 MiMo-V2.5Xiaomi
1,431
1,424–1,437
$0.17 / $0.34
95 Grok 4.1 ThinkingxAI
1,430
1,424–1,435
–
96 DeepSeek-V4-Flash (0423)DeepSeek
1,428
1,422–1,434
$0.14 / $0.28
96 MiMo-V2-OmniXiaomi
1,428
1,419–1,436
–
98 GPT-5.2 (high reasoning)OpenAI
1,427
1,421–1,433
$1.75 / $14
98 ChatGPT-4o (2025-03-26)OpenAI
1,427
1,422–1,431
–
100 GPT-5.1OpenAI
1,426
1,421–1,432
$1.25 / $10
100 Qwen3.8-27BAlibaba
1,426
1,418–1,434
$0.50 / $3
100 InklingThinking Machines
1,426
1,420–1,432
$0.95 / $4.05
103 GLM-4.7Z.ai
1,425
1,415–1,436
$0.54 / $1.98
103 ERNIE 5.0 (0110)Baidu
1,425
1,419–1,431
–
105 Qwen3-Max-PreviewAlibaba
1,424
1,417–1,432
–
105 GPT-5.2OpenAI
1,424
1,419–1,429
$1.75 / $14
105 GLM-5V-TurboZ.ai
1,424
1,414–1,434
$1.20 / $4
108 Mistral Medium 3.5Mistral AI
1,422
1,412–1,432
$1.50 / $7.50
109 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek
1,420
1,400–1,440
$0.27 / $1
109 DeepSeek-V3.2DeepSeek
1,420
1,414–1,426
$0.30 / $0.96
109 DeepSeek-V3.2 (reasoning on)DeepSeek
1,420
1,414–1,425
$0.30 / $0.96
112 ERNIE 5.0 Preview (1203)Baidu
1,419
1,408–1,431
–
113 Kimi K2 Thinking TurboMoonshot AI
1,418
1,413–1,423
–
114 DeepSeek-V3.1 (reasoning on)DeepSeek
1,416
1,405–1,427
$0.55 / $1.65
114 DeepSeek-V3.2-Exp (reasoning on)DeepSeek
1,416
1,404–1,427
$0.27 / $0.41
114 LongCat-Flash-Chat (2602, experimental)Meituan
1,416
1,409–1,423
–
117 Claude Sonnet 4 (32k reasoning budget)Anthropic
1,415
1,409–1,422
$3 / $15
117 Claude Haiku 4.5Anthropic
1,415
1,411–1,419
$1 / $5
117 Claude Opus 4Anthropic
1,415
1,409–1,422
–
117 GPT-5 ChatOpenAI
1,415
1,408–1,422
–
121 DeepSeek-V3.2-ExpDeepSeek
1,414
1,404–1,425
$0.27 / $0.41
121 Grok 4.3xAI
1,414
1,409–1,420
$1.25 / $2.50
121 Qwen3-235B-A22B-Instruct-2507Alibaba
1,414
1,410–1,418
$0.15 / $0.75
121 Nova Experimental Chat (2026-02-10)Amazon
1,414
1,396–1,432
–
125 Qwen3-MaxAlibaba
1,413
1,402–1,425
$0.78 / $3.90
125 Muse GlimmerMeta
1,413
1,397–1,429
–
127 GLM-4.6Z.ai
1,412
1,406–1,419
$0.50 / $2
127 Claude 3.7 Sonnet (32k reasoning budget)Anthropic
1,412
1,405–1,418
–
129 GPT-5 (high reasoning)OpenAI
1,410
1,403–1,417
$1.25 / $10
130 Qwen3-VL-235B-A22B-InstructAlibaba
1,409
1,398–1,421
$0.30 / $1.50
131 Gemini 3.1 Flash-Lite PreviewGoogle
1,407
1,402–1,413
$0.25 / $1.50
131 o1OpenAI
1,407
1,401–1,414
$15 / $60
131 MiniMax-M2.7MiniMax
1,407
1,402–1,412
$0.30 / $1.20
134 Qwen3.5-122B-A10BAlibaba
1,406
1,399–1,412
$0.26 / $2.08
135 Nemotron 3 Ultra 550B A55B (NVFP4)NVIDIA
1,405
1,395–1,415
–
135 GLM-4.5Z.ai
1,405
1,397–1,412
$0.60 / $2.20
137 GPT-4.1OpenAI
1,403
1,398–1,409
$2 / $8
137 Grok 3 Preview (02-24)xAI
1,403
1,397–1,410
–
137 o3OpenAI
1,403
1,397–1,408
$2 / $8
140 DeepSeek-V3.1DeepSeek
1,402
1,392–1,412
$0.55 / $1.65
140 Grok 4.1 Fast (reasoning)xAI
1,402
1,396–1,407
–
142 Qwen3.5-27BAlibaba
1,401
1,394–1,407
$0.27 / $2.16
143 Mistral Large 3Mistral AI
1,400
1,396–1,405
$0.50 / $1.50
143 Gemini 2.5 Flash Preview (09-2025)Google
1,400
1,393–1,407
–
145 Gemini 2.5 FlashGoogle
1,399
1,396–1,403
$0.30 / $2.50
146 Nova Experimental Chat (2026-01-10)Amazon
1,398
1,380–1,416
–
146 ERNIE 5.0 Preview (1022)Baidu
1,398
1,383–1,413
–
148 DeepSeek-R1DeepSeek
1,397
1,390–1,405
$0.70 / $2.50
148 Hunyuan Vision 1.5 ThinkingTencent
1,397
1,375–1,419
–
148 Grok 4 (0709)xAI
1,397
1,390–1,403
–
151 Hy3 PreviewTencent
1,396
1,384–1,409
$0.18 / $0.60
151 Claude Sonnet 4Anthropic
1,396
1,389–1,403
$3 / $15
153 Grok 4 Fast ChatxAI
1,395
1,381–1,410
–
153 Mistral Medium 3.1Mistral AI
1,395
1,391–1,399
$0.40 / $2
155 DeepSeek-V3.1-TerminusDeepSeek
1,394
1,377–1,412
$0.27 / $1
155 Inkling SmallThinking Machines
1,394
1,387–1,400
$0.45 / $1.20
157 DeepSeek-R1-0528DeepSeek
1,391
1,381–1,401
$0.50 / $2.18
158 Kimi K2 (0905)Moonshot AI
1,390
1,379–1,401
$0.60 / $2.50
159 Grok 4 Fast (reasoning)xAI
1,388
1,380–1,397
–
159 LongCat-Flash-ChatMeituan
1,388
1,377–1,399
–
161 GPT-5.4 nano (high reasoning)OpenAI
1,387
1,381–1,392
$0.20 / $1.25
162 Qwen3.5-35B-A3BAlibaba
1,386
1,379–1,392
$0.16 / $1.30
163 Claude 3.7 SonnetAnthropic
1,385
1,380–1,391
–
163 MiniMax-M2.1 PreviewMiniMax
1,385
1,376–1,394
–
163 Qwen3-235B-A22B-Thinking-2507Alibaba
1,385
1,372–1,397
$0.30 / $3
166 Step 3.5 FlashStepFun
1,384
1,379–1,389
$0.10 / $0.30
166 Qwen3-Coder-480B-A35BAlibaba
1,384
1,376–1,391
$0.35 / $1.50
168 Qwen3-VL-235B-A22B-ThinkingAlibaba
1,382
1,370–1,394
$0.40 / $4
168 MiMo-V2-Flash (no reasoning)Xiaomi
1,382
1,377–1,388
–
170 MiniMax-M2.5MiniMax
1,381
1,375–1,387
$0.30 / $1.20
170 Qwen3.5-FlashAlibaba
1,381
1,376–1,386
$0.065 / $0.26
170 o1-previewOpenAI
1,381
1,374–1,388
–
170 Qwen3-235B-A22B (no reasoning)Alibaba
1,381
1,374–1,388
$0.46 / $1.82
170 Kimi K2 (0711)Moonshot AI
1,381
1,373–1,388
$0.57 / $2.30
175 Solar Pro 4Upstage
1,379
1,371–1,387
$0.09 / $0.36
176 DeepSeek-V3-0324DeepSeek
1,378
1,372–1,384
$0.25 / $1
177 Nova Experimental Chat (12-10)Amazon
1,377
1,359–1,396
–
178 Qwen3-Next-80B-A3B-InstructAlibaba
1,376
1,369–1,384
$0.10 / $1.10
178 MiMo-V2-Flash (reasoning on)Xiaomi
1,376
1,365–1,386
–
180 Claude 3.5 Sonnet (2024-10-22)Anthropic
1,374
1,370–1,379
–
181 Hunyuan T1 (2025-07-11)Tencent
1,373
1,356–1,391
–
181 GPT-5 mini (high reasoning)OpenAI
1,373
1,366–1,381
$0.25 / $2
181 GPT-4.1 miniOpenAI
1,373
1,366–1,379
$0.40 / $1.60
184 Trinity Large PreviewArcee AI
1,370
1,363–1,376
–
185 o4-miniOpenAI
1,369
1,362–1,375
$1.10 / $4.40
186 Gemini 2.5 Flash-Lite Preview (06-17, reasoning on)Google
1,368
1,361–1,375
–
187 GLM-4.6VZ.ai
1,367
1,346–1,388
$0.30 / $0.90
187 Qwen3-30B-A3B-Instruct-2507Alibaba
1,367
1,359–1,374
$0.09 / $0.30
189 o3-mini (high reasoning)OpenAI
1,366
1,359–1,374
$1.10 / $4.40
189 Mistral Medium 3Mistral AI
1,366
1,359–1,373
$0.40 / $2
191 Gemini 2.5 Flash-Lite Preview (09-2025, no reasoning)Google
1,363
1,358–1,369
–
192 Nova Experimental Chat (11-10)Amazon
1,361
1,354–1,368
–
192 GLM-4.5-AirZ.ai
1,361
1,354–1,368
$0.14 / $0.86
194 Grok 3 Mini (high reasoning)xAI
1,360
1,351–1,369
–
195 Qwen3-Next-80B-A3B-ThinkingAlibaba
1,359
1,350–1,369
$0.15 / $1.20
196 Trinity Large ThinkingArcee AI
1,357
1,350–1,364
$0.25 / $0.80
196 Qwen3-235B-A22BAlibaba
1,357
1,349–1,364
$0.46 / $1.82
196 Qwen2.5-MaxAlibaba
1,357
1,351–1,363
–
199 Hunyuan TurboS (2025-04-16)Tencent
1,353
1,341–1,365
–
199 Grok 3 Mini BetaxAI
1,353
1,345–1,361
–
201 Llama 3.1 Nemotron Ultra 253B v1NVIDIA
1,351
1,330–1,373
–
201 Hunyuan TurboS (2025-02-26)Tencent
1,351
1,333–1,368
–
203 GLM-4.7-FlashZ.ai
1,348
1,338–1,358
$0.06 / $0.40
204 Gemini 2.0 Flash (001)Google
1,347
1,342–1,353
–
205 MiniMax-M1MiniMax
1,346
1,339–1,353
$0.40 / $2.20
206 Nova Experimental Chat (10-20)Amazon
1,345
1,335–1,356
–
206 Step 3StepFun
1,345
1,330–1,359
–
208 Nemotron 3.5 Lightning 30B A3B (NVFP4)NVIDIA
1,344
1,335–1,353
–
208 o3-miniOpenAI
1,344
1,339–1,349
$1.10 / $4.40
210 DeepSeek-V3DeepSeek
1,343
1,336–1,350
$0.26 / $1.03
210 Gemma 3 27BGoogle
1,343
1,337–1,349
$0.12 / $0.20
212 Nemotron 3 Super 120B A12BNVIDIA
1,342
1,330–1,355
$0.085 / $0.40
212 Command ACohere
1,342
1,336–1,347
$2.50 / $10
214 Gemini 1.5 Pro (002)Google
1,340
1,335–1,345
–
215 GLM-4.5VZ.ai
1,339
1,323–1,355
$0.60 / $1.80
216 Claude 3.5 Sonnet (2024-06-20)Anthropic
1,338
1,333–1,343
–
216 Granite 4.2 30BIBM
1,338
1,320–1,356
–
216 Mistral Small 3.2 24BMistral AI
1,338
1,329–1,347
$0.094 / $0.25
219 MiniMax-M2MiniMax
1,334
1,321–1,347
$0.30 / $1.20
220 INTELLECT-3Prime Intellect
1,332
1,317–1,348
–
220 Qwen-PlusAlibaba
1,332
1,320–1,344
$0.26 / $0.78
220 o1-miniOpenAI
1,332
1,326–1,337
–
223 Gemini 2.0 Flash-Lite Preview (02-05)Google
1,331
1,325–1,338
–
223 Qwen3-32BAlibaba
1,331
1,312–1,350
$0.14 / $0.40
225 Nova 2 LiteAmazon
1,327
1,317–1,337
$0.30 / $2.50
225 GPT-5 nano (high reasoning)OpenAI
1,327
1,314–1,340
$0.05 / $0.40
227 Mercury 2Inception
1,326
1,307–1,346
$0.25 / $0.75
227 GPT-4o (2024-05-13)OpenAI
1,326
1,321–1,331
$5 / $15
229 Llama 3.3 Nemotron Super 49B v1NVIDIA
1,325
1,306–1,345
–
230 gpt-oss-120bOpenAI
1,323
1,316–1,330
$0.15 / $0.60
230 QwQ-32BAlibaba
1,323
1,316–1,330
–
232 Llama 3.3 Nemotron Super 49B v1.5NVIDIA
1,321
1,301–1,341
–
232 Gemma 3 12BGoogle
1,321
1,305–1,337
$0.05 / $0.15
234 Olmo 3.1 32B InstructAi2
1,320
1,310–1,331
–
235 GPT-4o (2024-08-06)OpenAI
1,319
1,313–1,325
$2.50 / $10
236 Gemini Advanced (0514)Google
1,317
1,311–1,324
–
237 Ring-flash-2.0Ant Group
1,316
1,303–1,330
–
237 Claude 3.5 HaikuAnthropic
1,316
1,311–1,321
–
237 Hunyuan Large (2025-02-10)Tencent
1,316
1,300–1,332
–
237 Step-2 16K Exp (2024-12)StepFun
1,316
1,303–1,328
–
241 Ling-flash-2.0Ant Group
1,315
1,302–1,329
–
241 Llama 3.1 405B Instruct (FP8)Meta
1,315
1,310–1,320
–
243 Claude 3 OpusAnthropic
1,314
1,310–1,319
–
243 GLM-4-Plus (0111)Z.ai
1,314
1,302–1,327
–
243 Nova Experimental Chat (10-09)Amazon
1,314
1,293–1,336
–
243 Llama 3.1 405B Instruct (BF16)Meta
1,314
1,309–1,320
–
243 DeepSeek-V2.5-1210DeepSeek
1,314
1,303–1,326
–
243 Llama 4 MaverickMeta
1,314
1,307–1,320
$0.27 / $0.85
243 Hunyuan Turbo (0110)Tencent
1,314
1,295–1,332
–
250 Grok 2 (2024-08-13)xAI
1,312
1,307–1,317
–
250 Gemini 1.5 Pro (001)Google
1,312
1,306–1,317
–
252 Qwen3-30B-A3BAlibaba
1,310
1,302–1,317
$0.12 / $0.50
253 Yi-Lightning01.AI
1,308
1,301–1,315
–
254 Step-1o Turbo (2025-06)StepFun
1,307
1,294–1,319
–
255 Molmo 2 8BAi2
1,304
1,267–1,342
–
256 GPT-4 TurboOpenAI
1,303
1,298–1,309
$10 / $30
257 Athene-V2-ChatNexusflow
1,302
1,296–1,308
–
257 Qwen-Max (0919)Alibaba
1,302
1,294–1,310
–
259 GLM-4-PlusZ.ai
1,301
1,294–1,308
–
259 Magistral Medium 1.0Mistral AI
1,301
1,290–1,311
–
261 GPT-4.1 nanoOpenAI
1,300
1,288–1,313
$0.10 / $0.40
261 Mistral Large 2 (2407)Mistral AI
1,300
1,294–1,305
$2 / $6
263 Llama 4 ScoutMeta
1,299
1,292–1,306
$0.18 / $0.59
263 Olmo 3 32B ThinkAi2
1,299
1,283–1,315
–
265 GPT-4 Turbo Preview (1106)OpenAI
1,297
1,291–1,302
–
266 Mistral Large 2.1 (2411)Mistral AI
1,295
1,289–1,301
–
267 Qwen2.5-Plus (1127)Alibaba
1,294
1,285–1,304
–
267 Mistral Small 3.1 24BMistral AI
1,294
1,287–1,301
$0.35 / $0.56
269 GPT-4o mini (2024-07-18)OpenAI
1,293
1,288–1,298
$0.15 / $0.60
269 Llama 3.3 70B InstructMeta
1,293
1,288–1,298
$0.59 / $0.79
271 Qwen2.5-72B-InstructAlibaba
1,292
1,286–1,298
$0.36 / $0.40
271 DeepSeek-V2.5DeepSeek
1,292
1,285–1,298
–
273 Nemotron 3 Nano 30B A3BNVIDIA
1,291
1,282–1,300
$0.05 / $0.20
274 GPT-4 Turbo Preview (0125)OpenAI
1,290
1,285–1,296
–
274 Gemini 1.5 Flash (002)Google
1,290
1,284–1,296
–
276 Granite 4.1 8BIBM
1,287
1,270–1,304
–
277 Hunyuan Large VisionTencent
1,286
1,269–1,303
–
278 GPT-4 (0314)OpenAI
1,285
1,278–1,293
–
278 Granite 4.2 3BIBM
1,285
1,265–1,305
–
280 Grok 2 Mini (2024-08-13)xAI
1,284
1,278–1,289
–
281 Gemma 3n E4BGoogle
1,282
1,273–1,290
–
281 Hunyuan Standard (2025-02-10)Tencent
1,282
1,266–1,297
–
283 gpt-oss-20bOpenAI
1,281
1,269–1,294
$0.03 / $0.15
283 Granite 4.2 8BIBM
1,281
1,262–1,300
$0.06 / $0.25
285 Llama 3.1 Nemotron 70B InstructNVIDIA
1,280
1,269–1,291
–
285 Olmo 3.1 32B ThinkAi2
1,280
1,266–1,293
–
287 Athene 70B (0725)Nexusflow
1,279
1,271–1,286
–
288 GPT-4 (0613)OpenAI
1,277
1,271–1,283
–
288 Nova Pro 1.0Amazon
1,277
1,270–1,283
$0.80 / $3.20
290 Llama 3.1 70B InstructMeta
1,273
1,267–1,278
$0.40 / $0.40
291 Llama 3.1 Tülu 3 70BAi2
1,272
1,257–1,288
–
292 Gemma 2 27BGoogle
1,271
1,267–1,276
$0.65 / $0.65
293 MercuryInception
1,269
1,244–1,295
–
294 Claude 3 SonnetAnthropic
1,268
1,263–1,274
–
294 Granite H SmallIBM
1,268
1,252–1,283
–
296 Gemma 3 4BGoogle
1,267
1,251–1,284
$0.05 / $0.10
296 Jamba 1.5 LargeAI21 Labs
1,267
1,256–1,277
–
298 Gemini 1.5 Flash (001)Google
1,265
1,259–1,271
–
299 Qwen2.5-Coder-32B-InstructAlibaba
1,264
1,252–1,276
$0.66 / $1
300 Llama 3.1 Nemotron 51B InstructNVIDIA
1,263
1,248–1,277
–
301 Reka Core (2024-09-04)Reka AI
1,262
1,252–1,272
–
302 Nemotron-4 340B InstructNVIDIA
1,259
1,251–1,268
–
302 Llama 3 70B InstructMeta
1,259
1,253–1,264
–
304 GLM-4 (0520)Z.ai
1,256
1,246–1,267
–
305 Mistral Small 3Mistral AI
1,255
1,247–1,264
$0.05 / $0.08
306 Command R+ (08-2024)Cohere
1,254
1,245–1,263
$2.50 / $10
307 DeepSeek-Coder-V2DeepSeek
1,253
1,244–1,262
–
308 Gemma 2 9B SimPOPrinceton NLP
1,252
1,242–1,262
–
309 Reka Flash (2024-09-04)Reka AI
1,249
1,239–1,260
–
309 Aya Expanse 32BCohere
1,249
1,243–1,256
–
311 Phi-4Microsoft
1,245
1,239–1,252
$0.07 / $0.14
311 Claude 3 HaikuAnthropic
1,245
1,240–1,251
–
311 Gemma 2 9BGoogle
1,245
1,240–1,251
–
314 Nova Lite 1.0Amazon
1,244
1,237–1,251
$0.06 / $0.24
315 Hunyuan Standard 256KTencent
1,243
1,227–1,260
–
315 Qwen2-72B-InstructAlibaba
1,243
1,236–1,249
–
317 Command R+Cohere
1,241
1,235–1,247
–
318 Gemini 1.5 Flash-8B (001)Google
1,238
1,232–1,244
–
318 Mistral Large (2402)Mistral AI
1,238
1,231–1,245
–
320 Command R (08-2024)Cohere
1,236
1,227–1,245
$0.15 / $0.60
321 OLMo 2 32B Instruct (0325)Ai2
1,229
1,212–1,246
–
322 Gemini ProGoogle
1,220
1,204–1,237
–
323 Qwen1.5-110B-ChatAlibaba
1,218
1,210–1,226
–
324 GPT-3.5 Turbo (0125)OpenAI
1,216
1,209–1,222
–
325 Mixtral 8x22B InstructMistral AI
1,215
1,209–1,222
$2 / $6
325 Nova Micro 1.0Amazon
1,215
1,208–1,222
$0.035 / $0.14
327 Qwen1.5-72B-ChatAlibaba
1,212
1,205–1,220
–
328 Ministral 8B (2410)Mistral AI
1,211
1,198–1,224
–
328 Mistral MediumMistral AI
1,211
1,202–1,219
–
330 Gemini Pro (Dev API)Google
1,210
1,200–1,221
–
331 Llama 3.1 Tülu 3 8BAi2
1,208
1,193–1,224
–
332 Jamba 1.5 MiniAI21 Labs
1,205
1,194–1,215
–
332 Aya Expanse 8BCohere
1,205
1,195–1,214
–
334 Reka Flash 21B (2024-02-26, online)Reka AI
1,202
1,192–1,212
–
335 Command RCohere
1,200
1,193–1,207
–
336 GPT-3.5 Turbo (1106)OpenAI
1,198
1,186–1,210
–
337 Zephyr ORPO 141B A35B v0.1Hugging Face
1,194
1,178–1,209
–
338 Granite 3.1 8B InstructIBM
1,193
1,176–1,210
–
338 Reka Flash 21B (2024-02-26)Reka AI
1,193
1,184–1,201
–
338 Llama 3 8B InstructMeta
1,193
1,187–1,199
–
341 Llama 3.1 8B InstructMeta
1,191
1,185–1,197
$0.05 / $0.08
342 Yi-1.5-34B-Chat01.AI
1,188
1,180–1,196
–
342 DBRX Instruct PreviewDatabricks
1,188
1,179–1,196
–
344 Qwen1.5-32B-ChatAlibaba
1,185
1,176–1,193
–
345 Mixtral 8x7B Instruct v0.1Mistral AI
1,181
1,175–1,187
–
346 InternLM2.5 20B ChatShanghai AI Lab
1,178
1,168–1,188
–
346 Phi-3-medium-4k-instructMicrosoft
1,178
1,170–1,185
–
348 Granite 3.1 2B InstructIBM
1,173
1,157–1,190
–
349 Granite 3.0 8B InstructIBM
1,172
1,160–1,185
–
350 Tülu 2 DPO 70BAi2
1,171
1,157–1,186
–
350 Gemma 2 2BGoogle
1,171
1,165–1,177
–
352 Qwen1.5-14B-ChatAlibaba
1,168
1,158–1,178
–
353 WizardLM 70BMicrosoft
1,164
1,150–1,177
–
354 DeepSeek LLM 67B ChatDeepSeek
1,158
1,141–1,175
–
355 Gemma 1.1 7BGoogle
1,157
1,149–1,165
–
356 OpenChat 3.5 (0106)OpenChat
1,156
1,145–1,166
–
357 OpenChat 3.5OpenChat
1,155
1,140–1,169
–
358 Phi-3-small-8k-instructMicrosoft
1,154
1,146–1,163
–
358 Snowflake Arctic InstructSnowflake
1,154
1,145–1,162
–
358 OpenHermes 2.5 Mistral 7BTeknium
1,154
1,138–1,169
–
361 Yi-34B-Chat01.AI
1,152
1,143–1,162
–
362 Llama 3.2 3B InstructMeta
1,146
1,135–1,157
$0.05 / $0.33
362 Starling-LM-7B-betaNexusflow
1,146
1,136–1,156
–
362 QwQ-32B-PreviewAlibaba
1,146
1,130–1,161
–
365 Vicuna 33BLMSYS
1,140
1,131–1,149
–
366 Starling-LM-7B-alphaUC Berkeley
1,137
1,125–1,148
–
367 MPT-30B-ChatMosaicML
1,135
1,114–1,156
–
367 Llama 2 70B ChatMeta
1,135
1,127–1,143
–
367 Falcon 180B ChatTII
1,135
1,106–1,163
–
370 Granite 3.0 2B InstructIBM
1,132
1,120–1,145
–
371 Dolphin 2.2.1 Mistral 7BCognitive Computations
1,129
1,105–1,154
–
372 Llama 2 70B SteerLM ChatNVIDIA
1,125
1,106–1,143
–
372 Phi-3-mini-4k-instruct (June 2024)Microsoft
1,125
1,115–1,134
–
374 WizardLM 13BMicrosoft
1,124
1,110–1,138
–
374 Qwen1.5-7B-ChatAlibaba
1,124
1,110–1,138
–
374 Mistral 7B Instruct v0.2Mistral AI
1,124
1,114–1,133
–
377 Qwen-14B-ChatAlibaba
1,119
1,103–1,135
–
378 PaLM 2Google
1,117
1,104–1,131
–
378 Vicuna 13BLMSYS
1,117
1,108–1,127
–
380 Solar 10.7B Instruct v1.0Upstage
1,116
1,098–1,135
–
381 Phi-3-mini-4k-instructMicrosoft
1,114
1,105–1,122
–
382 Llama 2 13B ChatMeta
1,110
1,101–1,120
–
383 Nous Hermes 2 Mixtral 8x7B DPONous Research
1,108
1,091–1,124
–
384 Gemma 7BGoogle
1,105
1,092–1,118
–
384 Code Llama 34B InstructMeta
1,105
1,091–1,118
–
386 SmolLM2 1.7B InstructHugging Face
1,104
1,083–1,125
–
387 Gemma 1.1 2BGoogle
1,102
1,091–1,113
–
388 Phi-3-mini-128k-instructMicrosoft
1,100
1,090–1,111
–
389 Zephyr 7B AlphaHugging Face
1,099
1,075–1,123
–
390 Code Llama 70B InstructMeta
1,098
1,068–1,127
–
391 StripedHyena Nous 7BTogether AI
1,091
1,076–1,106
–
392 Zephyr 7B BetaHugging Face
1,089
1,076–1,102
–
393 Mistral 7B InstructMistral AI
1,087
1,074–1,101
–
394 Llama 3.2 1B InstructMeta
1,086
1,075–1,097
$0.027 / $0.20
395 Vicuna 7BLMSYS
1,077
1,063–1,091
–
396 Gemma 2BGoogle
1,071
1,055–1,087
–
397 Qwen1.5-4B-ChatAlibaba
1,070
1,057–1,083
–
397 Llama 2 7B ChatMeta
1,070
1,060–1,080
–
399 Guanaco 33BTim Dettmers
1,066
1,045–1,087
–
400 GPT4All 13B SnoozyNomic AI
1,038
1,013–1,063
–
401 ChatGLM3-6BZ.ai
1,037
1,019–1,055
–
402 OLMo 7B InstructAi2
1,029
1,013–1,045
–
403 Koala 13BUC Berkeley
1,027
1,011–1,042
–
404 Alpaca 13BStanford
1,025
1,009–1,042
–
405 MPT-7B-ChatMosaicML
1,009
991–1,028
–
406 OpenAssistant Pythia 12BOpenAssistant
988
972–1,004
–
407 ChatGLM2-6BZ.ai
983
960–1,005
–
408 ChatGLM-6BZ.ai
978
960–996
–
409 RWKV-4 Raven 14BRWKV
970
953–987
–
410 FastChat-T5 3BLMSYS
958
940–977
–
411 Dolly v2 12BDatabricks
940
919–961
–
412 LLaMA 13BMeta
918
893–944
–
413 StableLM Tuned Alpha 7BStability AI
910
890–930
–

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. ≈ marks results whose 95% range overlaps the leader's: they cannot be told apart from it. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Human preference on prompts that set explicit instructions.

What it does not measure

Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.