Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Human preference combined with labels of whether each answer's checkable claims were true (factuality weighted 25% by default).

Last updated 2 Oct 2026

Results dated
2 Oct 2026
Results
182 configurations of 163 models
Unit
Arena rating
Licence
Creative Commons Attribution 4.0 International

Overall: Gemini 4 Argon

Top 15 of 163 results · Arena rating, higher is better · lines show the 95% range · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ Gemini 4 Argon (high reasoning)Google 1,511
  2. 2≈ Claude Opus 5.5 (high reasoning)Anthropic 1,507
  3. 3≈ Claude Fable 5.1 (max reasoning)Anthropic 1,500
  4. 4 Claude Opus 4.6 (high reasoning)Anthropic 1,494
  5. 5 Claude Opus 5 (max reasoning)Anthropic 1,489
  6. 6 GPT-5.5 (high reasoning)OpenAI 1,486
  7. 7 GPT-5.4 (high reasoning)OpenAI 1,485
  8. 8 Claude Fable 5 (high reasoning)Anthropic 1,484
  9. 9 MiMo-V2.6-ProXiaomi 1,483
  10. 9 Muse Spark 1.2 (extra-high reasoning)Meta 1,483
  11. 11 Gemini 3 Pro PreviewGoogle 1,481
  12. 11 Qwen3.7-Max-PreviewAlibaba 1,481
  13. 13 Muse Spark 1.3 (max reasoning)Meta 1,480
  14. 14 Claude Opus 4.7 (high reasoning)Anthropic 1,477
  15. 14 Gemini 3.7 Flash (high reasoning)Google 1,477

Full results

Arena Text factuality: overall, Arena rating, higher is better
#ModelOverall · 95% range
Arena rating, higher is better
Price
$ per million tokens, in / out
1≈ Gemini 4 Argon (high reasoning)Google
1,511
1,503–1,518
–
2≈ Claude Opus 5.5 (high reasoning)Anthropic
1,507
1,499–1,515
$4 / $20
3≈ Claude Fable 5.1 (max reasoning)Anthropic
1,500
1,494–1,505
$10 / $50
4 Claude Opus 4.6 (high reasoning)Anthropic · best of 2 settings
1,494
1,491–1,496
$5 / $25
5 Claude Opus 5 (max reasoning)Anthropic · best of 2 settings
1,489
1,485–1,493
$5 / $25
6 GPT-5.5 (high reasoning)OpenAI · best of 2 settings
1,486
1,483–1,489
$5 / $30
7 GPT-5.4 (high reasoning)OpenAI · best of 2 settings
1,485
1,482–1,488
$2.50 / $15
8 Claude Fable 5 (high reasoning)Anthropic
1,484
1,480–1,487
$10 / $50
9 MiMo-V2.6-ProXiaomi
1,483
1,475–1,491
$0.43 / $0.87
9 Muse Spark 1.2 (extra-high reasoning)Meta
1,483
1,475–1,491
$1.25 / $4.25
11 Gemini 3 Pro PreviewGoogle
1,481
1,478–1,485
–
11 Qwen3.7-Max-PreviewAlibaba
1,481
1,472–1,489
–
13 Muse Spark 1.3 (max reasoning)Meta
1,480
1,475–1,485
$1.25 / $4.25
14 Claude Opus 4.7 (high reasoning)Anthropic · best of 2 settings
1,477
1,474–1,480
$5 / $25
14 Gemini 3.7 Flash (high reasoning)Google
1,477
1,473–1,481
$1.50 / $7.50
14 Gemini 3.8 Flash (high reasoning)Google
1,477
1,472–1,481
$1.50 / $7.50
17 Gemini 3.5 Flash (high reasoning)Google · best of 2 settings
1,476
1,473–1,480
$1.50 / $9
17 GPT-5.6 Sol (extra-high reasoning)OpenAI
1,476
1,473–1,480
$4 / $20
19 Qwen3.5-Max-PreviewAlibaba
1,475
1,470–1,479
–
20 Kimi K3 (max reasoning)Moonshot AI
1,472
1,469–1,476
$3 / $15
20 Gemini 3.1 Pro PreviewGoogle
1,472
1,470–1,474
$2 / $12
20 Qwen3.8-Max (0902)Alibaba
1,472
1,467–1,476
$2 / $6
23 Gemini 3.6 Flash (high reasoning)Google
1,471
1,468–1,475
$1.50 / $7.50
24 Claude Sonnet 5.5 (extra-high reasoning)Anthropic
1,470
1,461–1,479
$2 / $10
24 Claude Opus 4.8 (high reasoning)Anthropic · best of 2 settings
1,470
1,467–1,473
$5 / $25
26 Muse SparkMeta
1,468
1,462–1,473
–
26 Gemini 3 Flash PreviewGoogle · best of 2 settings
1,468
1,464–1,471
$0.50 / $3
28 DeepSeek-V4.1-Flash (max reasoning)DeepSeek
1,467
1,461–1,473
$0.30 / $1.20
28 GLM-5.3 (max reasoning)Z.ai
1,467
1,462–1,472
$1.40 / $4.40
30 GPT-5.6 Terra (extra-high reasoning)OpenAI
1,466
1,463–1,470
$2 / $12
30 GPT-6 Astra (max reasoning)OpenAI
1,466
1,460–1,472
$10 / $50
30 GLM-5.3-FlashZ.ai
1,466
1,461–1,470
$0.15 / $0.50
33 Grok 4.5xAI
1,465
1,462–1,469
$2 / $6
33 MiMo-V2.5-ProXiaomi
1,465
1,462–1,468
$0.43 / $0.87
35 Muse Spark 1.1Meta
1,464
1,460–1,467
$1.25 / $4.25
36 GLM-5.2 (max reasoning)Z.ai
1,463
1,460–1,467
$1.40 / $4.40
36 GPT-6.1 Sol (max reasoning)OpenAI
1,463
1,454–1,472
$2 / $10
38 Gemini 2.5 ProGoogle
1,462
1,459–1,465
$1.25 / $10
39 Claude Sonnet 4.6Anthropic
1,461
1,458–1,464
$3 / $15
40 GPT-5.6 Luna (extra-high reasoning)OpenAI
1,460
1,457–1,464
$0.20 / $1.20
41 Claude Opus 4.5Anthropic · best of 2 settings
1,459
1,456–1,462
$5 / $25
41 GPT-5.1 (high reasoning)OpenAI · best of 2 settings
1,459
1,455–1,462
$1.25 / $10
43 MiMo-V2.6-FlashXiaomi
1,457
1,450–1,464
$0.14 / $0.28
43 Grok 4.20 Multi-Agent Beta (0309)xAI
1,457
1,454–1,460
–
45 Kimi K2.6Moonshot AI
1,456
1,452–1,460
$0.95 / $4
45 GLM-5.1Z.ai
1,456
1,453–1,459
$1.38 / $4.40
45 ERNIE 5.1Baidu
1,456
1,452–1,460
–
45 Grok 4.20 Beta (0309, reasoning)xAI
1,456
1,453–1,459
–
49 DeepSeek-V4-Pro (0423)DeepSeek · best of 2 settings
1,455
1,452–1,458
$1.42 / $2.83
50 DeepSeek-V4-Pro (0813, high reasoning)DeepSeek
1,454
1,448–1,459
$1.32 / $3.96
51 Qwen3.6-Max-PreviewAlibaba
1,453
1,445–1,460
$1.03 / $6.16
51 Qwen3.7-PlusAlibaba
1,453
1,449–1,456
$0.32 / $1.28
53 Grok 4.20 Beta 1xAI
1,452
1,448–1,457
–
53 Step 5 PreviewStepFun
1,452
1,442–1,461
–
53 Claude Sonnet 5 (high reasoning)Anthropic
1,452
1,448–1,455
$2 / $10
53 GLM-5Z.ai
1,452
1,448–1,456
$0.95 / $2.55
57 Nova Experimental Chat (2026-02-10)Amazon
1,451
1,443–1,460
–
57 InklingThinking Machines
1,451
1,448–1,455
$0.95 / $4.05
57 GPT-5.2 Chat (2026-02-10)OpenAI
1,451
1,447–1,455
–
60 Gemma 4 31BGoogle
1,450
1,443–1,457
$0.14 / $0.40
61 Claude Sonnet 4.5Anthropic · best of 2 settings
1,449
1,446–1,451
$3 / $15
61 Grok 4.6 (high reasoning)xAI
1,449
1,444–1,453
$2 / $6
63 Gemini 3.5 Flash-LiteGoogle
1,447
1,444–1,451
$0.30 / $2.50
63 GPT-5.4 mini (high reasoning)OpenAI
1,447
1,444–1,450
$0.75 / $4.50
63 GLM-4.6Z.ai
1,447
1,443–1,451
$0.50 / $2
66 DeepSeek-V3.2-Exp (reasoning on)DeepSeek · best of 2 settings
1,446
1,435–1,457
$0.27 / $0.41
67 Hy3Tencent
1,445
1,440–1,451
$0.14 / $0.58
67 ERNIE 5.0 (0110)Baidu
1,445
1,442–1,449
–
67 Kimi K2.5 (reasoning on)Moonshot AI · best of 2 settings
1,445
1,442–1,448
$0.57 / $2.85
67 Qwen3.5-397B-A17BAlibaba
1,445
1,442–1,447
$0.55 / $3.50
67 Gemma 4 26B A4BGoogle
1,445
1,438–1,452
$0.10 / $0.30
72 GLM-4.7Z.ai
1,444
1,438–1,450
$0.54 / $1.98
72 ChatGPT-4o (2025-03-26)OpenAI
1,444
1,441–1,447
–
72 GPT-5.5 InstantOpenAI
1,444
1,439–1,448
–
75 MiMo-V2-ProXiaomi
1,443
1,439–1,448
–
76 GPT-5.2OpenAI · best of 2 settings
1,442
1,440–1,445
$1.75 / $14
76 ERNIE 5.0 Preview (1203)Baidu
1,442
1,436–1,448
–
76 Qwen3-Max-PreviewAlibaba
1,442
1,435–1,448
–
79 GLM-5V-TurboZ.ai
1,440
1,435–1,446
$1.20 / $4
79 ERNIE 5.0 Preview (1022)Baidu
1,440
1,431–1,449
–
81 DeepSeek-V4-Flash (0423)DeepSeek · best of 2 settings
1,439
1,436–1,442
$0.14 / $0.28
81 MiMo-V2.5Xiaomi
1,439
1,435–1,442
$0.17 / $0.34
83 Qwen3.8-27BAlibaba
1,438
1,434–1,443
$0.50 / $3
83 Qwen3.6-PlusAlibaba
1,438
1,434–1,441
$0.33 / $1.95
85 DeepSeek-V3.2DeepSeek · best of 2 settings
1,437
1,434–1,440
$0.30 / $0.96
85 Gemini 2.5 FlashGoogle
1,437
1,434–1,439
$0.30 / $2.50
87 MiMo-V2-OmniXiaomi
1,436
1,431–1,441
–
88 Claude Opus 4.1Anthropic · best of 2 settings
1,435
1,432–1,439
$15 / $75
89 MiniMax-M3MiniMax
1,434
1,431–1,437
$0.30 / $1.20
89 GPT-6 Sol (max reasoning)OpenAI
1,434
1,428–1,440
$2 / $10
89 Grok 4.1 ThinkingxAI
1,434
1,431–1,437
–
89 Mistral Large 3Mistral AI
1,434
1,431–1,436
$0.50 / $1.50
93 Gemini 3.1 Flash-Lite PreviewGoogle
1,433
1,430–1,436
$0.25 / $1.50
94 Qwen3-235B-A22B-Instruct-2507Alibaba
1,431
1,429–1,434
$0.15 / $0.75
94 Nemotron 3 Ultra 550B A55B (NVFP4)NVIDIA
1,431
1,425–1,437
–
94 Grok 4.1xAI
1,431
1,428–1,434
–
94 Mistral Medium 3.5Mistral AI
1,431
1,425–1,436
$1.50 / $7.50
98 GPT-6 Luna (max reasoning)OpenAI
1,430
1,424–1,436
$0.10 / $0.50
99 Mistral Medium 3.1Mistral AI
1,429
1,426–1,432
$0.40 / $2
99 LongCat-Flash-Chat (2602, experimental)Meituan
1,429
1,425–1,433
–
101 Grok 4.7 (extra-high reasoning)xAI
1,428
1,421–1,435
$2 / $6
101 Grok 4.3xAI
1,428
1,425–1,431
$1.25 / $2.50
101 Qwen3.5-122B-A10BAlibaba
1,428
1,424–1,431
$0.26 / $2.08
104 Inkling SmallThinking Machines
1,426
1,422–1,430
$0.45 / $1.20
105 Gemini 2.5 Flash Preview (09-2025)Google
1,425
1,420–1,429
–
106 Nova Experimental Chat (12-10)Amazon
1,424
1,415–1,434
–
107 Kimi K2 Thinking TurboMoonshot AI
1,423
1,420–1,426
–
108 MiMo-V2-Flash (no reasoning)Xiaomi · best of 2 settings
1,422
1,419–1,425
–
108 GPT-5 ChatOpenAI
1,422
1,415–1,429
–
108 MiniMax-M2.7MiniMax
1,422
1,419–1,425
$0.30 / $1.20
111 GPT-5 (high reasoning)OpenAI
1,421
1,414–1,428
$1.25 / $10
112 Qwen3.5-27BAlibaba
1,420
1,416–1,424
$0.27 / $2.16
113 Dola-Seed-2.0-ProByteDance
1,419
1,416–1,422
–
113 Qwen3-Next-80B-A3B-InstructAlibaba
1,419
1,412–1,425
$0.10 / $1.10
113 GPT-5.3 ChatOpenAI
1,419
1,415–1,422
–
116 Grok 4 (0709)xAI
1,418
1,412–1,424
–
116 GPT-5.4 nano (high reasoning)OpenAI
1,418
1,415–1,421
$0.20 / $1.25
116 Claude Haiku 4.5Anthropic
1,418
1,416–1,420
$1 / $5
119 Grok 4 Fast (reasoning)xAI
1,417
1,411–1,423
–
119 Step 3.5 FlashStepFun
1,417
1,414–1,420
$0.10 / $0.30
121 Nova Experimental Chat (11-10)Amazon
1,415
1,411–1,419
–
121 Qwen3-VL-235B-A22B-InstructAlibaba
1,415
1,397–1,432
$0.30 / $1.50
123 Qwen3.5-FlashAlibaba
1,414
1,411–1,417
$0.065 / $0.26
124 Grok 4.1 Fast (reasoning)xAI
1,413
1,410–1,416
–
125 Hy3 PreviewTencent
1,412
1,405–1,419
$0.18 / $0.60
126 MiniMax-M2.1 PreviewMiniMax
1,411
1,406–1,416
–
127 Solar Pro 4Upstage
1,410
1,405–1,415
$0.09 / $0.36
128 GPT-4.1OpenAI
1,408
1,402–1,415
$2 / $8
129 Qwen3.5-35B-A3BAlibaba
1,407
1,404–1,411
$0.16 / $1.30
129 Gemini 2.5 Flash-Lite Preview (09-2025, no reasoning)Google
1,407
1,404–1,411
–
131 o3OpenAI
1,405
1,398–1,411
$2 / $8
132 Kimi K2 (0905)Moonshot AI
1,403
1,389–1,418
$0.60 / $2.50
132 Nova Experimental Chat (2026-01-10)Amazon
1,403
1,394–1,411
–
134 Muse GlimmerMeta
1,402
1,394–1,410
–
135 Qwen3-Coder-480B-A35BAlibaba
1,401
1,383–1,419
$0.35 / $1.50
136 Nova Experimental Chat (10-20)Amazon
1,399
1,393–1,405
–
137 GLM-4.5-AirZ.ai
1,395
1,388–1,402
$0.14 / $0.86
138 GPT-5 mini (high reasoning)OpenAI
1,394
1,387–1,401
$0.25 / $2
139 GLM-4.6VZ.ai
1,392
1,381–1,402
$0.30 / $0.90
140 Nemotron 3 Super 120B A12BNVIDIA
1,390
1,383–1,397
$0.085 / $0.40
141 MiniMax-M2.5MiniMax
1,386
1,382–1,389
$0.30 / $1.20
142 gpt-oss-120bOpenAI
1,385
1,378–1,391
$0.15 / $0.60
143 Nova 2 LiteAmazon
1,380
1,374–1,386
$0.30 / $2.50
143 Gemma 3 27BGoogle
1,380
1,369–1,391
$0.12 / $0.20
145 Granite 4.2 30BIBM
1,375
1,366–1,384
–
146 INTELLECT-3Prime Intellect
1,373
1,365–1,381
–
146 Trinity Large PreviewArcee AI
1,373
1,369–1,376
–
148 Trinity Large ThinkingArcee AI
1,371
1,367–1,375
$0.25 / $0.80
149 GLM-4.7-FlashZ.ai
1,370
1,364–1,376
$0.06 / $0.40
149 MiniMax-M1MiniMax
1,370
1,359–1,381
$0.40 / $2.20
151 Mercury 2Inception
1,369
1,359–1,378
$0.25 / $0.75
152 o4-miniOpenAI
1,368
1,358–1,378
$1.10 / $4.40
153 Command ACohere
1,363
1,357–1,370
$2.50 / $10
154 Nemotron 3.5 Lightning 30B A3B (NVFP4)NVIDIA
1,362
1,357–1,368
–
154 Nemotron 3 Nano 30B A3BNVIDIA
1,362
1,357–1,367
$0.05 / $0.20
156 MiniMax-M2MiniMax
1,360
1,353–1,367
$0.30 / $1.20
157 Olmo 3.1 32B InstructAi2
1,342
1,336–1,348
–
157 Granite 4.2 8BIBM
1,342
1,332–1,352
$0.06 / $0.25
159 Granite 4.2 3BIBM
1,325
1,315–1,335
–
160 Olmo 3 32B ThinkAi2
1,319
1,312–1,327
–
161 Granite 4.1 8BIBM
1,311
1,302–1,320
–
162 MercuryInception
1,305
1,292–1,317
–
163 Olmo 3.1 32B ThinkAi2
1,300
1,293–1,307
–

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. ≈ marks results whose 95% range overlaps the leader's: they cannot be told apart from it. Each model is shown at its best setting; show every setting. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Human preference combined with labels of whether each answer's checkable claims were true (factuality weighted 25% by default).

What it does not measure

Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.