Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Head-to-head human preference across all text prompts.

Last updated 2 Oct 2026

Results dated
2 Oct 2026
Results
413 configurations of 387 models
Unit
Arena rating
Licence
Creative Commons Attribution 4.0 International

Overall: Gemini 3.7 Flash

Top 15 of 413 results · Arena rating, higher is better · lines show the 95% range. Choose a model to highlight it.Clear highlight

  1. 1 Gemini 4 Argon (high reasoning)Google 1,525
  2. 2 Claude Opus 4.6 (high reasoning)Anthropic 1,505
  3. 3 Claude Fable 5 (high reasoning)Anthropic 1,504
  4. 3 Claude Opus 5.5 (high reasoning)Anthropic 1,504
  5. 5 Claude Opus 4.7 (high reasoning)Anthropic 1,501
  6. 5 Claude Fable 5.1 (max reasoning)Anthropic 1,501
  7. 7 Claude Opus 4.6Anthropic 1,497
  8. 8 Gemini 3.8 Flash (high reasoning)Google 1,495
  9. 9 Muse Spark 1.3 (max reasoning)Meta 1,494
  10. 9 Claude Opus 4.7Anthropic 1,494
  11. 9 Muse Spark 1.2 (extra-high reasoning)Meta 1,494
  12. 12 Muse Spark 1.1Meta 1,491
  13. 13 Claude Opus 5 (high reasoning)Anthropic 1,490
  14. 14 Claude Opus 5 (max reasoning)Anthropic 1,489
  15. 14 Muse SparkMeta 1,489
  16. 16 Gemini 3.7 Flash (high reasoning)Google 1,488

Full results

Arena Text: overall, Arena rating, higher is better
#ModelOverall · 95% range
Arena rating, higher is better
Price
$ per million tokens, in / out
1 Gemini 4 Argon (high reasoning)Google
1,525
1,516–1,534
–
2 Claude Opus 4.6 (high reasoning)Anthropic
1,505
1,501–1,508
$5 / $25
3 Claude Fable 5 (high reasoning)Anthropic
1,504
1,500–1,509
$10 / $50
3 Claude Opus 5.5 (high reasoning)Anthropic
1,504
1,495–1,513
$4 / $20
5 Claude Opus 4.7 (high reasoning)Anthropic
1,501
1,498–1,505
$5 / $25
5 Claude Fable 5.1 (max reasoning)Anthropic
1,501
1,495–1,508
$10 / $50
7 Claude Opus 4.6Anthropic
1,497
1,494–1,501
$5 / $25
8 Gemini 3.8 Flash (high reasoning)Google
1,495
1,490–1,500
$1.50 / $7.50
9 Muse Spark 1.3 (max reasoning)Meta
1,494
1,488–1,500
$1.25 / $4.25
9 Claude Opus 4.7Anthropic
1,494
1,490–1,498
$5 / $25
9 Muse Spark 1.2 (extra-high reasoning)Meta
1,494
1,484–1,503
$1.25 / $4.25
12 Muse Spark 1.1Meta
1,491
1,487–1,496
$1.25 / $4.25
13 Claude Opus 5 (high reasoning)Anthropic
1,490
1,486–1,494
$5 / $25
14 Claude Opus 5 (max reasoning)Anthropic
1,489
1,485–1,494
$5 / $25
14 Muse SparkMeta
1,489
1,484–1,495
–
16 Kimi K3 (max reasoning)Moonshot AI
1,488
1,484–1,493
$3 / $15
16 Gemini 3.7 Flash (high reasoning)Google
1,488
1,483–1,493
$1.50 / $7.50
18 Gemini 3.1 Pro PreviewGoogle
1,487
1,484–1,490
$2 / $12
19 Gemini 3 Pro PreviewGoogle
1,485
1,482–1,489
–
20 GPT-5.6 Sol (extra-high reasoning)OpenAI
1,484
1,479–1,488
$4 / $20
21 GPT-6.1 Sol (max reasoning)OpenAI
1,483
1,473–1,494
$2 / $10
21 Gemini 3.6 Flash (high reasoning)Google
1,483
1,479–1,487
$1.50 / $7.50
23 Qwen3.8-Max (0902)Alibaba
1,482
1,476–1,487
$2 / $6
23 Claude Opus 4.8 (high reasoning)Anthropic
1,482
1,478–1,485
$5 / $25
25 GPT-5.5 (high reasoning)OpenAI
1,481
1,478–1,485
$5 / $30
26 MiMo-V2.6-ProXiaomi
1,480
1,471–1,489
$0.43 / $0.87
27 GLM-5.3 (max reasoning)Z.ai
1,478
1,473–1,484
$1.40 / $4.40
28 Gemini 3.5 Flash (high reasoning)Google
1,477
1,473–1,481
$1.50 / $9
28 GPT-6 Astra (max reasoning)OpenAI
1,477
1,470–1,484
$10 / $50
28 GPT-5.5OpenAI
1,477
1,473–1,480
$5 / $30
31 Gemini 3.5 Flash (medium reasoning)Google
1,476
1,472–1,480
$1.50 / $9
31 GLM-5.2 (max reasoning)Z.ai
1,476
1,471–1,480
$1.40 / $4.40
31 GPT-5.2 Chat (2026-02-10)OpenAI
1,476
1,472–1,480
–
34 GPT-5.4 (high reasoning)OpenAI
1,475
1,471–1,479
$2.50 / $15
34 Qwen3.7-Max-PreviewAlibaba
1,475
1,465–1,485
–
34 Grok 4.20 Beta 1xAI
1,475
1,470–1,479
–
34 Claude Opus 4.8Anthropic
1,475
1,471–1,478
$5 / $25
38 DeepSeek-V4.1-Flash (max reasoning)DeepSeek
1,474
1,468–1,481
$0.30 / $1.20
38 Claude Opus 4.5 (high reasoning, 32k budget)Anthropic
1,474
1,470–1,477
$5 / $25
40 GPT-5.5 InstantOpenAI
1,473
1,468–1,478
–
40 GLM-5.3-FlashZ.ai
1,473
1,468–1,478
$0.15 / $0.50
40 Gemini 3 Flash PreviewGoogle
1,473
1,468–1,477
$0.50 / $3
43 Claude Sonnet 4.6Anthropic
1,472
1,469–1,476
$3 / $15
43 Grok 4.20 Beta (0309, reasoning)xAI
1,472
1,468–1,475
–
45 Claude Sonnet 5.5 (extra-high reasoning)Anthropic
1,471
1,461–1,481
$2 / $10
45 Grok 4.20 Multi-Agent Beta (0309)xAI
1,471
1,467–1,474
–
47 Claude Opus 4.5Anthropic
1,470
1,467–1,473
$5 / $25
48 ERNIE 5.1Baidu
1,468
1,463–1,472
–
48 MiMo-V2.5-ProXiaomi
1,468
1,464–1,471
$0.43 / $0.87
50 Grok 4.5xAI
1,466
1,461–1,470
$2 / $6
51 GPT-5.6 Terra (extra-high reasoning)OpenAI
1,465
1,461–1,470
$2 / $12
51 Grok 4.1 ThinkingxAI
1,465
1,462–1,468
–
51 GPT-5.4OpenAI
1,465
1,461–1,468
$2.50 / $15
51 Qwen3.5-Max-PreviewAlibaba
1,465
1,460–1,470
–
51 GLM-5.1Z.ai
1,465
1,461–1,468
$1.38 / $4.40
56 DeepSeek-V4-Pro (0813, high reasoning)DeepSeek
1,464
1,458–1,471
$1.32 / $3.96
57 Claude Sonnet 5 (high reasoning)Anthropic
1,462
1,458–1,466
$2 / $10
58 Kimi K2.6Moonshot AI
1,461
1,457–1,465
$0.95 / $4
59 Qwen3.6-Max-PreviewAlibaba
1,460
1,452–1,469
$1.03 / $6.16
60 Grok 4.1xAI
1,459
1,456–1,462
–
61 Gemini 3 Flash Preview (minimal reasoning)Google
1,458
1,455–1,461
$0.50 / $3
61 DeepSeek-V4-Pro (0423)DeepSeek
1,458
1,454–1,462
$1.42 / $2.83
61 GLM-5Z.ai
1,458
1,453–1,462
$0.95 / $2.55
64 GPT-6 Sol (max reasoning)OpenAI
1,457
1,450–1,464
$2 / $10
64 Claude Sonnet 4.5 (high reasoning, 32k budget)Anthropic
1,457
1,454–1,459
$3 / $15
66 Dola-Seed-2.0-ProByteDance
1,456
1,453–1,460
–
66 Hy3Tencent
1,456
1,449–1,463
$0.14 / $0.58
66 Qwen3.7-PlusAlibaba
1,456
1,451–1,460
$0.32 / $1.28
66 GPT-5.1 (high reasoning)OpenAI
1,456
1,452–1,459
$1.25 / $10
70 Claude Sonnet 4.5Anthropic
1,455
1,452–1,458
$3 / $15
70 Gemini 3.5 Flash-LiteGoogle
1,455
1,450–1,459
$0.30 / $2.50
70 DeepSeek-V4-Pro (0423, high reasoning)DeepSeek
1,455
1,451–1,459
$1.42 / $2.83
73 Grok 4.6 (high reasoning)xAI
1,454
1,449–1,459
$2 / $6
73 Step 5 PreviewStepFun
1,454
1,443–1,465
–
75 GPT-5.6 Luna (extra-high reasoning)OpenAI
1,453
1,449–1,457
$0.20 / $1.20
75 Gemma 4 31BGoogle
1,453
1,445–1,460
$0.14 / $0.40
77 MiMo-V2.6-FlashXiaomi
1,452
1,444–1,460
$0.14 / $0.28
78 Kimi K2.5 (reasoning on)Moonshot AI
1,450
1,447–1,453
$0.57 / $2.85
78 GPT-5.3 ChatOpenAI
1,450
1,446–1,454
–
78 Claude Opus 4.1 (16k reasoning budget)Anthropic
1,450
1,446–1,453
$15 / $75
81 ERNIE 5.0 Preview (1203)Baidu
1,449
1,442–1,455
–
82 Claude Opus 4.1Anthropic
1,448
1,445–1,451
$15 / $75
82 MiMo-V2-ProXiaomi
1,448
1,443–1,453
–
84 GPT-5.4 mini (high reasoning)OpenAI
1,447
1,443–1,451
$0.75 / $4.50
85 ERNIE 5.0 (0110)Baidu
1,446
1,442–1,450
–
85 Gemini 2.5 ProGoogle
1,446
1,443–1,448
$1.25 / $10
87 GPT-4.5 PreviewOpenAI
1,445
1,439–1,450
–
88 Qwen3.6-PlusAlibaba
1,443
1,439–1,448
$0.33 / $1.95
88 ChatGPT-4o (2025-03-26)OpenAI
1,443
1,440–1,446
–
88 GPT-6 Luna (max reasoning)OpenAI
1,443
1,436–1,449
$0.10 / $0.50
91 Grok 4.7 (extra-high reasoning)xAI
1,442
1,435–1,450
$2 / $6
91 Qwen3.5-397B-A17BAlibaba
1,442
1,439–1,445
$0.55 / $3.50
93 GLM-4.7Z.ai
1,441
1,435–1,448
$0.54 / $1.98
93 Grok 4.3xAI
1,441
1,438–1,445
$1.25 / $2.50
93 InklingThinking Machines
1,441
1,437–1,446
$0.95 / $4.05
96 MiniMax-M3MiniMax
1,440
1,436–1,444
$0.30 / $1.20
97 GPT-5.1OpenAI
1,439
1,435–1,442
$1.25 / $10
98 DeepSeek-V4-Flash (0423, high reasoning)DeepSeek
1,438
1,434–1,443
$0.14 / $0.28
98 Qwen3.8-27BAlibaba
1,438
1,432–1,443
$0.50 / $3
100 Gemma 4 26B A4BGoogle
1,437
1,430–1,445
$0.10 / $0.30
100 GPT-5.2 (high reasoning)OpenAI
1,437
1,434–1,441
$1.75 / $14
100 LongCat-Flash-Chat (2602, experimental)Meituan
1,437
1,432–1,441
–
103 DeepSeek-V4-Flash (0423)DeepSeek
1,436
1,432–1,440
$0.14 / $0.28
103 GPT-5.2OpenAI
1,436
1,432–1,439
$1.75 / $14
105 GPT-5 (high reasoning)OpenAI
1,435
1,430–1,439
$1.25 / $10
106 Qwen3-Max-PreviewAlibaba
1,434
1,430–1,439
–
106 MiMo-V2.5Xiaomi
1,434
1,429–1,438
$0.17 / $0.34
108 GLM-5V-TurboZ.ai
1,433
1,427–1,440
$1.20 / $4
108 Gemini 3.1 Flash-Lite PreviewGoogle
1,433
1,429–1,436
$0.25 / $1.50
110 o3OpenAI
1,432
1,428–1,435
$2 / $8
111 MiMo-V2-OmniXiaomi
1,431
1,425–1,437
–
111 Kimi K2.5 (no reasoning)Moonshot AI
1,431
1,424–1,437
$0.57 / $2.85
113 Grok 4.1 Fast (reasoning)xAI
1,430
1,427–1,434
–
113 Kimi K2 Thinking TurboMoonshot AI
1,430
1,427–1,433
–
115 Mistral Medium 3.5Mistral AI
1,427
1,421–1,433
$1.50 / $7.50
115 GPT-5 ChatOpenAI
1,427
1,422–1,431
–
115 Nova Experimental Chat (2026-02-10)Amazon
1,427
1,417–1,436
–
118 Nemotron 3 Ultra 550B A55B (NVFP4)NVIDIA
1,426
1,419–1,433
–
118 Claude Opus 4 (16k reasoning budget)Anthropic
1,426
1,422–1,430
–
120 Muse GlimmerMeta
1,425
1,415–1,435
–
120 DeepSeek-V3.2-Exp (reasoning on)DeepSeek
1,425
1,418–1,432
$0.27 / $0.41
120 DeepSeek-V3.2DeepSeek
1,425
1,421–1,428
$0.30 / $0.96
123 GLM-4.6Z.ai
1,424
1,420–1,428
$0.50 / $2
124 Qwen3-MaxAlibaba
1,423
1,416–1,429
$0.78 / $3.90
124 DeepSeek-V3.2 (reasoning on)DeepSeek
1,423
1,419–1,426
$0.30 / $0.96
126 Qwen3-235B-A22B-Instruct-2507Alibaba
1,422
1,420–1,425
$0.15 / $0.75
126 DeepSeek-V3.2-ExpDeepSeek
1,422
1,416–1,428
$0.27 / $0.41
126 DeepSeek-R1-0528DeepSeek
1,422
1,416–1,427
$0.50 / $2.18
129 Grok 4 Fast ChatxAI
1,419
1,411–1,427
–
129 ERNIE 5.0 Preview (1022)Baidu
1,419
1,410–1,428
–
129 Kimi K2 (0905)Moonshot AI
1,419
1,412–1,425
$0.60 / $2.50
132 Kimi K2 (0711)Moonshot AI
1,418
1,413–1,423
$0.57 / $2.30
133 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek
1,417
1,407–1,427
$0.27 / $1
133 DeepSeek-V3.1DeepSeek
1,417
1,411–1,423
$0.55 / $1.65
135 Qwen3.5-122B-A10BAlibaba
1,416
1,412–1,420
$0.26 / $2.08
135 DeepSeek-V3.1 (reasoning on)DeepSeek
1,416
1,409–1,423
$0.55 / $1.65
137 MiniMax-M2.7MiniMax
1,415
1,412–1,418
$0.30 / $1.20
137 DeepSeek-V3.1-TerminusDeepSeek
1,415
1,405–1,425
$0.27 / $1
137 GPT-4.1OpenAI
1,415
1,411–1,418
$2 / $8
140 Claude Opus 4Anthropic
1,414
1,410–1,418
–
140 Mistral Large 3Mistral AI
1,414
1,411–1,417
$0.50 / $1.50
140 Claude Haiku 4.5Anthropic
1,414
1,411–1,416
$1 / $5
143 Qwen3-VL-235B-A22B-InstructAlibaba
1,413
1,407–1,420
$0.30 / $1.50
143 Nova Experimental Chat (2026-01-10)Amazon
1,413
1,403–1,423
–
145 Grok 3 Preview (02-24)xAI
1,411
1,407–1,416
–
145 GLM-4.5Z.ai
1,411
1,406–1,416
$0.60 / $2.20
145 Grok 4 (0709)xAI
1,411
1,407–1,415
–
148 Gemini 2.5 FlashGoogle
1,409
1,407–1,412
$0.30 / $2.50
148 Hy3 PreviewTencent
1,409
1,401–1,416
$0.18 / $0.60
148 Qwen3.5-27BAlibaba
1,409
1,404–1,413
$0.27 / $2.16
151 Mistral Medium 3.1Mistral AI
1,408
1,405–1,411
$0.40 / $2
152 Inkling SmallThinking Machines
1,405
1,400–1,410
$0.45 / $1.20
152 Grok 4 Fast (reasoning)xAI
1,405
1,400–1,410
–
154 Gemini 2.5 Flash Preview (09-2025)Google
1,403
1,399–1,407
–
154 Qwen3-235B-A22B (no reasoning)Alibaba
1,403
1,398–1,407
$0.46 / $1.82
156 o1OpenAI
1,402
1,398–1,407
$15 / $60
156 Claude Sonnet 4 (32k reasoning budget)Anthropic
1,402
1,397–1,406
$3 / $15
158 GPT-5.4 nano (high reasoning)OpenAI
1,401
1,398–1,405
$0.20 / $1.25
158 LongCat-Flash-ChatMeituan
1,401
1,395–1,408
–
160 Qwen3-235B-A22B-Thinking-2507Alibaba
1,400
1,393–1,407
$0.30 / $3
161 Qwen3-Next-80B-A3B-InstructAlibaba
1,399
1,394–1,404
$0.10 / $1.10
162 DeepSeek-R1DeepSeek
1,398
1,393–1,403
$0.70 / $2.50
163 Qwen3.5-FlashAlibaba
1,396
1,392–1,400
$0.065 / $0.26
163 DeepSeek-V3-0324DeepSeek
1,396
1,392–1,400
$0.25 / $1
165 Qwen3-VL-235B-A22B-ThinkingAlibaba
1,395
1,388–1,402
$0.40 / $4
165 Hunyuan Vision 1.5 ThinkingTencent
1,395
1,382–1,407
–
167 Nova Experimental Chat (12-10)Amazon
1,394
1,384–1,403
–
167 Qwen3.5-35B-A3BAlibaba
1,394
1,390–1,398
$0.16 / $1.30
169 Step 3.5 FlashStepFun
1,393
1,390–1,397
$0.10 / $0.30
170 MiMo-V2-Flash (no reasoning)Xiaomi
1,392
1,388–1,395
–
171 MiniMax-M2.5MiniMax
1,391
1,387–1,395
$0.30 / $1.20
171 o4-miniOpenAI
1,391
1,387–1,395
$1.10 / $4.40
171 Claude Sonnet 4Anthropic
1,391
1,386–1,395
$3 / $15
174 GPT-5 mini (high reasoning)OpenAI
1,390
1,385–1,394
$0.25 / $2
175 o1-previewOpenAI
1,389
1,384–1,394
–
176 Claude 3.7 Sonnet (32k reasoning budget)Anthropic
1,388
1,384–1,393
–
176 Mistral Medium 3Mistral AI
1,388
1,383–1,392
$0.40 / $2
178 Qwen3-Coder-480B-A35BAlibaba
1,387
1,382–1,392
$0.35 / $1.50
178 Hunyuan T1 (2025-07-11)Tencent
1,387
1,378–1,396
–
178 MiMo-V2-Flash (reasoning on)Xiaomi
1,387
1,381–1,393
–
181 Solar Pro 4Upstage
1,386
1,380–1,392
$0.09 / $0.36
182 MiniMax-M2.1 PreviewMiniMax
1,385
1,379–1,390
–
183 GPT-4.1 miniOpenAI
1,383
1,378–1,387
$0.40 / $1.60
183 Hunyuan TurboS (2025-04-16)Tencent
1,383
1,376–1,389
–
183 Qwen3-30B-A3B-Instruct-2507Alibaba
1,383
1,378–1,388
$0.09 / $0.30
186 Gemini 2.5 Flash-Lite Preview (09-2025, no reasoning)Google
1,379
1,376–1,383
–
186 GLM-4.6VZ.ai
1,379
1,367–1,390
$0.30 / $0.90
186 Trinity Large PreviewArcee AI
1,379
1,374–1,383
–
189 Qwen3-235B-A22BAlibaba
1,375
1,371–1,380
$0.46 / $1.82
189 Gemini 2.5 Flash-Lite Preview (06-17, reasoning on)Google
1,375
1,370–1,379
–
191 Claude 3.5 Sonnet (2024-10-22)Anthropic
1,374
1,371–1,377
–
191 Qwen2.5-MaxAlibaba
1,374
1,370–1,378
–
193 GLM-4.5-AirZ.ai
1,373
1,369–1,378
$0.14 / $0.86
193 Claude 3.7 SonnetAnthropic
1,373
1,369–1,377
–
195 Qwen3-Next-80B-A3B-ThinkingAlibaba
1,369
1,363–1,375
$0.15 / $1.20
196 Trinity Large ThinkingArcee AI
1,367
1,363–1,372
$0.25 / $0.80
197 Gemma 3 27BGoogle
1,365
1,362–1,369
$0.12 / $0.20
197 GLM-4.7-FlashZ.ai
1,365
1,359–1,371
$0.06 / $0.40
199 Nova Experimental Chat (11-10)Amazon
1,364
1,360–1,369
–
199 MiniMax-M1MiniMax
1,364
1,360–1,368
$0.40 / $2.20
199 o3-mini (high reasoning)OpenAI
1,364
1,358–1,369
$1.10 / $4.40
202 Grok 3 Mini (high reasoning)xAI
1,363
1,358–1,369
–
203 Nemotron 3 Super 120B A12BNVIDIA
1,360
1,353–1,368
$0.085 / $0.40
203 Gemini 2.0 Flash (001)Google
1,360
1,356–1,364
–
205 DeepSeek-V3DeepSeek
1,358
1,354–1,363
$0.26 / $1.03
205 Grok 3 Mini BetaxAI
1,358
1,353–1,363
–
207 Mistral Small 3.2 24BMistral AI
1,356
1,351–1,362
$0.094 / $0.25
207 INTELLECT-3Prime Intellect
1,356
1,348–1,364
–
209 Command ACohere
1,354
1,350–1,357
$2.50 / $10
209 Gemini 2.0 Flash-Lite Preview (02-05)Google
1,354
1,349–1,358
–
211 GLM-4.5VZ.ai
1,352
1,344–1,361
$0.60 / $1.80
211 gpt-oss-120bOpenAI
1,352
1,347–1,356
$0.15 / $0.60
213 Gemini 1.5 Pro (002)Google
1,351
1,348–1,355
–
214 Nemotron 3.5 Lightning 30B A3B (NVFP4)NVIDIA
1,350
1,343–1,356
–
215 Step 3StepFun
1,349
1,342–1,357
–
215 Hunyuan TurboS (2025-02-26)Tencent
1,349
1,337–1,361
–
217 Nova Experimental Chat (10-20)Amazon
1,348
1,342–1,354
–
217 o3-miniOpenAI
1,348
1,345–1,352
$1.10 / $4.40
217 Llama 3.1 Nemotron Ultra 253B v1NVIDIA
1,348
1,336–1,359
–
220 Nova Experimental Chat (10-09)Amazon
1,347
1,336–1,358
–
220 Qwen3-32BAlibaba
1,347
1,337–1,356
$0.14 / $0.40
222 GPT-4o (2024-05-13)OpenAI
1,346
1,343–1,350
$5 / $15
222 Qwen-PlusAlibaba
1,346
1,338–1,354
$0.26 / $0.78
224 Ling-flash-2.0Ant Group
1,344
1,336–1,351
–
225 MiniMax-M2MiniMax
1,343
1,336–1,351
$0.30 / $1.20
225 Mercury 2Inception
1,343
1,333–1,354
$0.25 / $0.75
225 Claude 3.5 Sonnet (2024-06-20)Anthropic
1,343
1,340–1,347
–
225 Llama 3.3 Nemotron Super 49B v1.5NVIDIA
1,343
1,333–1,353
–
225 GLM-4-Plus (0111)Z.ai
1,343
1,334–1,351
–
230 Gemma 3 12BGoogle
1,342
1,332–1,351
$0.05 / $0.15
231 Hunyuan Turbo (0110)Tencent
1,341
1,329–1,352
–
232 Granite 4.2 30BIBM
1,340
1,329–1,350
–
233 GPT-5 nano (high reasoning)OpenAI
1,338
1,331–1,344
$0.05 / $0.40
234 o1-miniOpenAI
1,337
1,333–1,340
–
235 Gemini Advanced (0514)Google
1,336
1,331–1,341
–
235 QwQ-32BAlibaba
1,336
1,331–1,340
–
235 Grok 2 (2024-08-13)xAI
1,336
1,332–1,339
–
235 GPT-4o (2024-08-06)OpenAI
1,336
1,331–1,340
$2.50 / $10
239 Llama 3.1 405B Instruct (BF16)Meta
1,335
1,332–1,339
–
239 Nova 2 LiteAmazon
1,335
1,329–1,341
$0.30 / $2.50
241 Step-2 16K Exp (2024-12)StepFun
1,334
1,325–1,342
–
241 Llama 3.1 405B Instruct (FP8)Meta
1,334
1,330–1,337
–
243 Olmo 3.1 32B InstructAi2
1,329
1,323–1,335
–
244 Yi-Lightning01.AI
1,328
1,323–1,333
–
244 Llama 3.3 Nemotron Super 49B v1NVIDIA
1,328
1,316–1,340
–
246 Llama 4 MaverickMeta
1,327
1,323–1,331
$0.27 / $0.85
247 Qwen3-30B-A3BAlibaba
1,326
1,322–1,331
$0.12 / $0.50
247 Hunyuan Large (2025-02-10)Tencent
1,326
1,317–1,336
–
249 Claude 3.5 HaikuAnthropic
1,325
1,321–1,328
–
250 GPT-4 TurboOpenAI
1,324
1,321–1,328
$10 / $30
250 Gemini 1.5 Pro (001)Google
1,324
1,320–1,328
–
252 DeepSeek-V2.5-1210DeepSeek
1,323
1,315–1,332
–
252 Ring-flash-2.0Ant Group
1,323
1,315–1,330
–
254 GPT-4.1 nanoOpenAI
1,322
1,315–1,330
$0.10 / $0.40
254 Claude 3 OpusAnthropic
1,322
1,319–1,325
–
254 Molmo 2 8BAi2
1,322
1,301–1,343
–
254 Llama 4 ScoutMeta
1,322
1,317–1,326
$0.18 / $0.59
258 Step-1o Turbo (2025-06)StepFun
1,320
1,313–1,327
–
259 GLM-4-PlusZ.ai
1,319
1,314–1,324
–
260 Qwen-Max (0919)Alibaba
1,318
1,312–1,324
–
260 Llama 3.3 70B InstructMeta
1,318
1,314–1,321
$0.59 / $0.79
260 GPT-4o mini (2024-07-18)OpenAI
1,318
1,314–1,321
$0.15 / $0.60
260 gpt-oss-20bOpenAI
1,318
1,311–1,324
$0.03 / $0.15
264 Gemma 3n E4BGoogle
1,317
1,312–1,323
–
265 Qwen2.5-Plus (1127)Alibaba
1,315
1,308–1,321
–
266 Mistral Large 2 (2407)Mistral AI
1,314
1,310–1,318
$2 / $6
266 Athene-V2-ChatNexusflow
1,314
1,310–1,319
–
266 Nemotron 3 Nano 30B A3BNVIDIA
1,314
1,308–1,319
$0.05 / $0.20
269 GPT-4 Turbo Preview (0125)OpenAI
1,313
1,309–1,317
–
269 GPT-4 Turbo Preview (1106)OpenAI
1,313
1,309–1,317
–
271 Hunyuan Standard (2025-02-10)Tencent
1,311
1,302–1,321
–
272 Gemini 1.5 Flash (002)Google
1,309
1,305–1,313
–
273 Grok 2 Mini (2024-08-13)xAI
1,308
1,305–1,312
–
274 DeepSeek-V2.5DeepSeek
1,307
1,302–1,312
–
274 Athene 70B (0725)Nexusflow
1,307
1,301–1,312
–
276 Olmo 3 32B ThinkAi2
1,306
1,298–1,315
–
276 Mistral Large 2.1 (2411)Mistral AI
1,306
1,301–1,310
–
278 MercuryInception
1,305
1,291–1,319
–
278 Magistral Medium 1.0Mistral AI
1,305
1,298–1,311
–
280 Granite 4.1 8BIBM
1,304
1,294–1,314
–
281 Gemma 3 4BGoogle
1,303
1,294–1,313
$0.05 / $0.10
281 Mistral Small 3.1 24BMistral AI
1,303
1,299–1,308
$0.35 / $0.56
281 Qwen2.5-72B-InstructAlibaba
1,303
1,299–1,307
$0.36 / $0.40
284 Llama 3.1 Nemotron 70B InstructNVIDIA
1,299
1,291–1,306
–
285 Hunyuan Large VisionTencent
1,294
1,285–1,303
–
286 Llama 3.1 70B InstructMeta
1,293
1,290–1,297
$0.40 / $0.40
287 Nova Pro 1.0Amazon
1,290
1,286–1,295
$0.80 / $3.20
287 Granite 4.2 3BIBM
1,290
1,278–1,301
–
289 Gemma 2 27BGoogle
1,289
1,286–1,293
$0.65 / $0.65
289 Jamba 1.5 LargeAI21 Labs
1,289
1,282–1,297
–
291 Granite 4.2 8BIBM
1,288
1,277–1,299
$0.06 / $0.25
291 Reka Core (2024-09-04)Reka AI
1,288
1,281–1,295
–
291 GPT-4 (0314)OpenAI
1,288
1,283–1,293
–
294 Granite H SmallIBM
1,287
1,279–1,295
–
294 Llama 3.1 Nemotron 51B InstructNVIDIA
1,287
1,277–1,297
–
296 Gemini 1.5 Flash (001)Google
1,286
1,282–1,291
–
296 Olmo 3.1 32B ThinkAi2
1,286
1,279–1,294
–
296 Llama 3.1 Tülu 3 70BAi2
1,286
1,276–1,297
–
299 Claude 3 SonnetAnthropic
1,281
1,277–1,285
–
300 Gemma 2 9B SimPOPrinceton NLP
1,280
1,273–1,287
–
301 Nemotron-4 340B InstructNVIDIA
1,277
1,272–1,282
–
301 Llama 3 70B InstructMeta
1,277
1,273–1,280
–
303 Command R+ (08-2024)Cohere
1,276
1,270–1,283
$2.50 / $10
303 GPT-4 (0613)OpenAI
1,276
1,272–1,280
–
305 Mistral Small 3Mistral AI
1,274
1,268–1,280
$0.05 / $0.08
306 GLM-4 (0520)Z.ai
1,273
1,266–1,280
–
307 Reka Flash (2024-09-04)Reka AI
1,272
1,265–1,279
–
308 Qwen2.5-Coder-32B-InstructAlibaba
1,271
1,262–1,279
$0.66 / $1
309 Aya Expanse 32BCohere
1,267
1,262–1,272
–
309 Gemma 2 9BGoogle
1,267
1,263–1,271
–
311 DeepSeek-Coder-V2DeepSeek
1,265
1,259–1,272
–
312 Qwen2-72B-InstructAlibaba
1,262
1,257–1,267
–
312 Command R+Cohere
1,262
1,257–1,266
–
312 Claude 3 HaikuAnthropic
1,262
1,258–1,265
–
315 Nova Lite 1.0Amazon
1,260
1,255–1,266
$0.06 / $0.24
316 Gemini 1.5 Flash-8B (001)Google
1,259
1,254–1,263
–
317 Phi-4Microsoft
1,256
1,252–1,261
$0.07 / $0.14
318 OLMo 2 32B Instruct (0325)Ai2
1,252
1,241–1,262
–
319 Command R (08-2024)Cohere
1,250
1,244–1,257
$0.15 / $0.60
320 Mistral Large (2402)Mistral AI
1,242
1,238–1,247
–
321 Nova Micro 1.0Amazon
1,241
1,236–1,246
$0.035 / $0.14
322 Jamba 1.5 MiniAI21 Labs
1,240
1,233–1,247
–
323 Ministral 8B (2410)Mistral AI
1,238
1,228–1,247
–
324 Gemini Pro (Dev API)Google
1,237
1,229–1,244
–
325 Qwen1.5-110B-ChatAlibaba
1,234
1,229–1,240
–
326 Reka Flash 21B (2024-02-26, online)Reka AI
1,233
1,226–1,241
–
326 Qwen1.5-72B-ChatAlibaba
1,233
1,228–1,239
–
326 Hunyuan Standard 256KTencent
1,233
1,221–1,245
–
329 Mixtral 8x22B InstructMistral AI
1,230
1,225–1,234
$2 / $6
330 Command RCohere
1,227
1,222–1,232
–
330 Reka Flash 21B (2024-02-26)Reka AI
1,227
1,221–1,233
–
332 GPT-3.5 Turbo (0125)OpenAI
1,226
1,221–1,230
–
333 Llama 3 8B InstructMeta
1,224
1,220–1,227
–
333 Gemini ProGoogle
1,224
1,212–1,235
–
335 Aya Expanse 8BCohere
1,223
1,216–1,230
–
335 Mistral MediumMistral AI
1,223
1,217–1,228
–
337 Llama 3.1 Tülu 3 8BAi2
1,220
1,210–1,231
–
338 Zephyr ORPO 141B A35B v0.1Hugging Face
1,213
1,202–1,224
–
338 Yi-1.5-34B-Chat01.AI
1,213
1,208–1,218
–
340 Llama 3.1 8B InstructMeta
1,211
1,207–1,215
$0.05 / $0.08
341 Granite 3.1 8B InstructIBM
1,208
1,197–1,220
–
342 GPT-3.5 Turbo (1106)OpenAI
1,204
1,195–1,213
–
342 Qwen1.5-32B-ChatAlibaba
1,204
1,198–1,210
–
344 Gemma 2 2BGoogle
1,200
1,196–1,204
–
345 Phi-3-medium-4k-instructMicrosoft
1,198
1,193–1,203
–
346 Mixtral 8x7B Instruct v0.1Mistral AI
1,197
1,193–1,201
–
347 DBRX Instruct PreviewDatabricks
1,195
1,189–1,202
–
348 Qwen1.5-14B-ChatAlibaba
1,191
1,184–1,198
–
348 InternLM2.5 20B ChatShanghai AI Lab
1,191
1,184–1,198
–
350 DeepSeek LLM 67B ChatDeepSeek
1,185
1,173–1,196
–
350 WizardLM 70BMicrosoft
1,185
1,175–1,194
–
352 Yi-34B-Chat01.AI
1,184
1,177–1,191
–
353 Granite 3.0 8B InstructIBM
1,183
1,174–1,192
–
353 OpenChat 3.5OpenChat
1,183
1,173–1,193
–
353 OpenChat 3.5 (0106)OpenChat
1,183
1,175–1,191
–
353 Gemma 1.1 7BGoogle
1,183
1,176–1,189
–
357 Snowflake Arctic InstructSnowflake
1,180
1,174–1,186
–
358 Granite 3.1 2B InstructIBM
1,179
1,168–1,190
–
359 Tülu 2 DPO 70BAi2
1,178
1,168–1,187
–
360 OpenHermes 2.5 Mistral 7BTeknium
1,176
1,166–1,186
–
361 Vicuna 33BLMSYS
1,173
1,167–1,179
–
362 Phi-3-small-8k-instructMicrosoft
1,171
1,165–1,177
–
362 Starling-LM-7B-betaNexusflow
1,171
1,164–1,178
–
362 Llama 2 70B ChatMeta
1,171
1,165–1,176
–
365 Starling-LM-7B-alphaUC Berkeley
1,167
1,159–1,175
–
365 Llama 3.2 3B InstructMeta
1,167
1,159–1,174
$0.05 / $0.33
367 Nous Hermes 2 Mixtral 8x7B DPONous Research
1,164
1,152–1,176
–
368 Granite 3.0 2B InstructIBM
1,157
1,148–1,165
–
369 Llama 2 70B SteerLM ChatNVIDIA
1,155
1,142–1,167
–
370 QwQ-32B-PreviewAlibaba
1,154
1,143–1,166
–
371 Solar 10.7B Instruct v1.0Upstage
1,152
1,139–1,166
–
371 Dolphin 2.2.1 Mistral 7BCognitive Computations
1,152
1,137–1,167
–
373 MPT-30B-ChatMosaicML
1,151
1,138–1,163
–
374 WizardLM 13BMicrosoft
1,149
1,140–1,159
–
374 Mistral 7B Instruct v0.2Mistral AI
1,149
1,143–1,156
–
376 Falcon 180B ChatTII
1,148
1,131–1,165
–
377 Qwen1.5-7B-ChatAlibaba
1,144
1,134–1,154
–
378 Phi-3-mini-4k-instruct (June 2024)Microsoft
1,143
1,137–1,150
–
379 Vicuna 13BLMSYS
1,142
1,135–1,148
–
380 Llama 2 13B ChatMeta
1,141
1,135–1,148
–
381 Qwen-14B-ChatAlibaba
1,139
1,128–1,150
–
381 PaLM 2Google
1,139
1,130–1,148
–
383 Gemma 7BGoogle
1,138
1,128–1,147
–
384 Code Llama 34B InstructMeta
1,137
1,128–1,146
–
385 Zephyr 7B BetaHugging Face
1,131
1,122–1,140
–
386 Phi-3-mini-128k-instructMicrosoft
1,130
1,122–1,137
–
387 Phi-3-mini-4k-instructMicrosoft
1,128
1,122–1,134
–
388 Guanaco 33BTim Dettmers
1,127
1,115–1,139
–
388 Zephyr 7B AlphaHugging Face
1,127
1,111–1,142
–
390 StripedHyena Nous 7BTogether AI
1,122
1,111–1,133
–
391 Code Llama 70B InstructMeta
1,119
1,101–1,137
–
392 Gemma 1.1 2BGoogle
1,117
1,109–1,124
–
393 Vicuna 7BLMSYS
1,115
1,106–1,124
–
393 SmolLM2 1.7B InstructHugging Face
1,115
1,100–1,129
–
395 Llama 3.2 1B InstructMeta
1,111
1,103–1,119
$0.027 / $0.20
396 Mistral 7B InstructMistral AI
1,110
1,101–1,120
–
397 Llama 2 7B ChatMeta
1,108
1,101–1,115
–
398 Gemma 2BGoogle
1,094
1,082–1,105
–
399 Qwen1.5-4B-ChatAlibaba
1,091
1,082–1,100
–
400 OLMo 7B InstructAi2
1,074
1,062–1,085
–
401 Koala 13BUC Berkeley
1,071
1,061–1,081
–
402 Alpaca 13BStanford
1,070
1,059–1,082
–
403 GPT4All 13B SnoozyNomic AI
1,068
1,052–1,083
–
404 MPT-7B-ChatMosaicML
1,063
1,051–1,075
–
405 ChatGLM3-6BZ.ai
1,057
1,045–1,068
–
406 RWKV-4 Raven 14BRWKV
1,042
1,031–1,054
–
407 ChatGLM2-6BZ.ai
1,025
1,011–1,038
–
408 OpenAssistant Pythia 12BOpenAssistant
1,023
1,013–1,034
–
409 ChatGLM-6BZ.ai
996
983–1,008
–
410 FastChat-T5 3BLMSYS
992
980–1,005
–
411 Dolly v2 12BDatabricks
982
969–996
–
412 LLaMA 13BMeta
975
959–991
–
413 StableLM Tuned Alpha 7BStability AI
953
941–966
–

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Where two models' ranges overlap, the gap between them may not be real. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Head-to-head human preference across all text prompts.

What it does not measure

Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.