Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Human preference combined with labels of whether each answer's checkable claims were true (factuality weighted 25% by default).

Last updated 2 Oct 2026

Results dated
2 Oct 2026
Results
182 configurations of 163 models
Unit
Arena rating
Licence
Creative Commons Attribution 4.0 International

Overall: Claude Opus 5

Top 16 of 182 results · Arena rating, higher is better · lines show the 95% range · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ Gemini 4 Argon (high reasoning)Google 1,511
  2. 2≈ Claude Opus 5.5 (high reasoning)Anthropic 1,507
  3. 3≈ Claude Fable 5.1 (max reasoning)Anthropic 1,500
  4. 4 Claude Opus 4.6 (high reasoning)Anthropic 1,494
  5. 5 Claude Opus 5 (max reasoning)Anthropic 1,489
  6. 6 Claude Opus 4.6Anthropic 1,488
  7. 7 GPT-5.5 (high reasoning)OpenAI 1,486
  8. 8 GPT-5.4 (high reasoning)OpenAI 1,485
  9. 9 Claude Fable 5 (high reasoning)Anthropic 1,484
  10. 10 MiMo-V2.6-ProXiaomi 1,483
  11. 10 Muse Spark 1.2 (extra-high reasoning)Meta 1,483
  12. 12 GPT-5.5OpenAI 1,482
  13. 13 Gemini 3 Pro PreviewGoogle 1,481
  14. 13 Qwen3.7-Max-PreviewAlibaba 1,481
  15. 15 Muse Spark 1.3 (max reasoning)Meta 1,480
  16. 16 Claude Opus 5 (high reasoning)Anthropic 1,477

Full results

Arena Text factuality: overall, Arena rating, higher is better
#ModelOverall · 95% range
Arena rating, higher is better
Price
$ per million tokens, in / out
1≈ Gemini 4 Argon (high reasoning)Google
1,511
1,503–1,518
–
2≈ Claude Opus 5.5 (high reasoning)Anthropic
1,507
1,499–1,515
$4 / $20
3≈ Claude Fable 5.1 (max reasoning)Anthropic
1,500
1,494–1,505
$10 / $50
4 Claude Opus 4.6 (high reasoning)Anthropic
1,494
1,491–1,496
$5 / $25
5 Claude Opus 5 (max reasoning)Anthropic
1,489
1,485–1,493
$5 / $25
6 Claude Opus 4.6Anthropic
1,488
1,486–1,491
$5 / $25
7 GPT-5.5 (high reasoning)OpenAI
1,486
1,483–1,489
$5 / $30
8 GPT-5.4 (high reasoning)OpenAI
1,485
1,482–1,488
$2.50 / $15
9 Claude Fable 5 (high reasoning)Anthropic
1,484
1,480–1,487
$10 / $50
10 MiMo-V2.6-ProXiaomi
1,483
1,475–1,491
$0.43 / $0.87
10 Muse Spark 1.2 (extra-high reasoning)Meta
1,483
1,475–1,491
$1.25 / $4.25
12 GPT-5.5OpenAI
1,482
1,479–1,485
$5 / $30
13 Gemini 3 Pro PreviewGoogle
1,481
1,478–1,485
–
13 Qwen3.7-Max-PreviewAlibaba
1,481
1,472–1,489
–
15 Muse Spark 1.3 (max reasoning)Meta
1,480
1,475–1,485
$1.25 / $4.25
16 Claude Opus 5 (high reasoning)Anthropic
1,477
1,474–1,481
$5 / $25
16 Claude Opus 4.7 (high reasoning)Anthropic
1,477
1,474–1,480
$5 / $25
16 Gemini 3.7 Flash (high reasoning)Google
1,477
1,473–1,481
$1.50 / $7.50
16 Gemini 3.8 Flash (high reasoning)Google
1,477
1,472–1,481
$1.50 / $7.50
20 Gemini 3.5 Flash (high reasoning)Google
1,476
1,473–1,480
$1.50 / $9
20 GPT-5.6 Sol (extra-high reasoning)OpenAI
1,476
1,473–1,480
$4 / $20
22 GPT-5.4OpenAI
1,475
1,472–1,478
$2.50 / $15
22 Claude Opus 4.7Anthropic
1,475
1,472–1,478
$5 / $25
22 Qwen3.5-Max-PreviewAlibaba
1,475
1,470–1,479
–
25 Kimi K3 (max reasoning)Moonshot AI
1,472
1,469–1,476
$3 / $15
25 Gemini 3.1 Pro PreviewGoogle
1,472
1,470–1,474
$2 / $12
25 Qwen3.8-Max (0902)Alibaba
1,472
1,467–1,476
$2 / $6
28 Gemini 3.6 Flash (high reasoning)Google
1,471
1,468–1,475
$1.50 / $7.50
29 Claude Sonnet 5.5 (extra-high reasoning)Anthropic
1,470
1,461–1,479
$2 / $10
29 Claude Opus 4.8 (high reasoning)Anthropic
1,470
1,467–1,473
$5 / $25
31 Muse SparkMeta
1,468
1,462–1,473
–
31 Gemini 3 Flash PreviewGoogle
1,468
1,464–1,471
$0.50 / $3
33 DeepSeek-V4.1-Flash (max reasoning)DeepSeek
1,467
1,461–1,473
$0.30 / $1.20
33 GLM-5.3 (max reasoning)Z.ai
1,467
1,462–1,472
$1.40 / $4.40
33 Gemini 3.5 Flash (medium reasoning)Google
1,467
1,463–1,470
$1.50 / $9
36 GPT-5.6 Terra (extra-high reasoning)OpenAI
1,466
1,463–1,470
$2 / $12
36 GPT-6 Astra (max reasoning)OpenAI
1,466
1,460–1,472
$10 / $50
36 GLM-5.3-FlashZ.ai
1,466
1,461–1,470
$0.15 / $0.50
39 Grok 4.5xAI
1,465
1,462–1,469
$2 / $6
39 MiMo-V2.5-ProXiaomi
1,465
1,462–1,468
$0.43 / $0.87
41 Claude Opus 4.8Anthropic
1,464
1,461–1,468
$5 / $25
41 Muse Spark 1.1Meta
1,464
1,460–1,467
$1.25 / $4.25
43 GLM-5.2 (max reasoning)Z.ai
1,463
1,460–1,467
$1.40 / $4.40
43 GPT-6.1 Sol (max reasoning)OpenAI
1,463
1,454–1,472
$2 / $10
45 Gemini 2.5 ProGoogle
1,462
1,459–1,465
$1.25 / $10
46 Claude Sonnet 4.6Anthropic
1,461
1,458–1,464
$3 / $15
47 GPT-5.6 Luna (extra-high reasoning)OpenAI
1,460
1,457–1,464
$0.20 / $1.20
48 Claude Opus 4.5Anthropic
1,459
1,456–1,462
$5 / $25
48 Claude Opus 4.5 (high reasoning, 32k budget)Anthropic
1,459
1,456–1,463
$5 / $25
48 GPT-5.1 (high reasoning)OpenAI
1,459
1,455–1,462
$1.25 / $10
51 MiMo-V2.6-FlashXiaomi
1,457
1,450–1,464
$0.14 / $0.28
51 Grok 4.20 Multi-Agent Beta (0309)xAI
1,457
1,454–1,460
–
53 Kimi K2.6Moonshot AI
1,456
1,452–1,460
$0.95 / $4
53 GLM-5.1Z.ai
1,456
1,453–1,459
$1.38 / $4.40
53 ERNIE 5.1Baidu
1,456
1,452–1,460
–
53 Grok 4.20 Beta (0309, reasoning)xAI
1,456
1,453–1,459
–
57 DeepSeek-V4-Pro (0423)DeepSeek
1,455
1,452–1,458
$1.42 / $2.83
58 DeepSeek-V4-Pro (0813, high reasoning)DeepSeek
1,454
1,448–1,459
$1.32 / $3.96
59 Qwen3.6-Max-PreviewAlibaba
1,453
1,445–1,460
$1.03 / $6.16
59 Qwen3.7-PlusAlibaba
1,453
1,449–1,456
$0.32 / $1.28
61 Grok 4.20 Beta 1xAI
1,452
1,448–1,457
–
61 Step 5 PreviewStepFun
1,452
1,442–1,461
–
61 Claude Sonnet 5 (high reasoning)Anthropic
1,452
1,448–1,455
$2 / $10
61 GLM-5Z.ai
1,452
1,448–1,456
$0.95 / $2.55
65 Nova Experimental Chat (2026-02-10)Amazon
1,451
1,443–1,460
–
65 InklingThinking Machines
1,451
1,448–1,455
$0.95 / $4.05
65 Gemini 3 Flash Preview (minimal reasoning)Google
1,451
1,449–1,454
$0.50 / $3
65 GPT-5.2 Chat (2026-02-10)OpenAI
1,451
1,447–1,455
–
69 Gemma 4 31BGoogle
1,450
1,443–1,457
$0.14 / $0.40
70 Claude Sonnet 4.5Anthropic
1,449
1,446–1,451
$3 / $15
70 Grok 4.6 (high reasoning)xAI
1,449
1,444–1,453
$2 / $6
72 DeepSeek-V4-Pro (0423, high reasoning)DeepSeek
1,448
1,444–1,451
$1.42 / $2.83
73 Gemini 3.5 Flash-LiteGoogle
1,447
1,444–1,451
$0.30 / $2.50
73 GPT-5.4 mini (high reasoning)OpenAI
1,447
1,444–1,450
$0.75 / $4.50
73 GLM-4.6Z.ai
1,447
1,443–1,451
$0.50 / $2
73 Claude Sonnet 4.5 (high reasoning, 32k budget)Anthropic
1,447
1,444–1,449
$3 / $15
77 DeepSeek-V3.2-Exp (reasoning on)DeepSeek
1,446
1,435–1,457
$0.27 / $0.41
78 Hy3Tencent
1,445
1,440–1,451
$0.14 / $0.58
78 ERNIE 5.0 (0110)Baidu
1,445
1,442–1,449
–
78 Kimi K2.5 (reasoning on)Moonshot AI
1,445
1,442–1,448
$0.57 / $2.85
78 Qwen3.5-397B-A17BAlibaba
1,445
1,442–1,447
$0.55 / $3.50
78 Gemma 4 26B A4BGoogle
1,445
1,438–1,452
$0.10 / $0.30
78 GPT-5.1OpenAI
1,445
1,441–1,448
$1.25 / $10
84 GLM-4.7Z.ai
1,444
1,438–1,450
$0.54 / $1.98
84 ChatGPT-4o (2025-03-26)OpenAI
1,444
1,441–1,447
–
84 GPT-5.5 InstantOpenAI
1,444
1,439–1,448
–
87 MiMo-V2-ProXiaomi
1,443
1,439–1,448
–
88 GPT-5.2OpenAI
1,442
1,440–1,445
$1.75 / $14
88 ERNIE 5.0 Preview (1203)Baidu
1,442
1,436–1,448
–
88 Qwen3-Max-PreviewAlibaba
1,442
1,435–1,448
–
88 GPT-5.2 (high reasoning)OpenAI
1,442
1,438–1,445
$1.75 / $14
92 GLM-5V-TurboZ.ai
1,440
1,435–1,446
$1.20 / $4
92 ERNIE 5.0 Preview (1022)Baidu
1,440
1,431–1,449
–
94 DeepSeek-V4-Flash (0423)DeepSeek
1,439
1,436–1,442
$0.14 / $0.28
94 MiMo-V2.5Xiaomi
1,439
1,435–1,442
$0.17 / $0.34
96 Qwen3.8-27BAlibaba
1,438
1,434–1,443
$0.50 / $3
96 Qwen3.6-PlusAlibaba
1,438
1,434–1,441
$0.33 / $1.95
98 DeepSeek-V3.2DeepSeek
1,437
1,434–1,440
$0.30 / $0.96
98 Gemini 2.5 FlashGoogle
1,437
1,434–1,439
$0.30 / $2.50
100 DeepSeek-V3.2-ExpDeepSeek
1,436
1,430–1,442
$0.27 / $0.41
100 MiMo-V2-OmniXiaomi
1,436
1,431–1,441
–
102 Claude Opus 4.1Anthropic
1,435
1,432–1,439
$15 / $75
103 MiniMax-M3MiniMax
1,434
1,431–1,437
$0.30 / $1.20
103 GPT-6 Sol (max reasoning)OpenAI
1,434
1,428–1,440
$2 / $10
103 Grok 4.1 ThinkingxAI
1,434
1,431–1,437
–
103 Mistral Large 3Mistral AI
1,434
1,431–1,436
$0.50 / $1.50
103 DeepSeek-V4-Flash (0423, high reasoning)DeepSeek
1,434
1,430–1,437
$0.14 / $0.28
108 Claude Opus 4.1 (16k reasoning budget)Anthropic
1,433
1,429–1,437
$15 / $75
108 Gemini 3.1 Flash-Lite PreviewGoogle
1,433
1,430–1,436
$0.25 / $1.50
108 DeepSeek-V3.2 (reasoning on)DeepSeek
1,433
1,429–1,436
$0.30 / $0.96
111 Qwen3-235B-A22B-Instruct-2507Alibaba
1,431
1,429–1,434
$0.15 / $0.75
111 Nemotron 3 Ultra 550B A55B (NVFP4)NVIDIA
1,431
1,425–1,437
–
111 Grok 4.1xAI
1,431
1,428–1,434
–
111 Mistral Medium 3.5Mistral AI
1,431
1,425–1,436
$1.50 / $7.50
115 GPT-6 Luna (max reasoning)OpenAI
1,430
1,424–1,436
$0.10 / $0.50
116 Kimi K2.5 (no reasoning)Moonshot AI
1,429
1,423–1,435
$0.57 / $2.85
116 Mistral Medium 3.1Mistral AI
1,429
1,426–1,432
$0.40 / $2
116 LongCat-Flash-Chat (2602, experimental)Meituan
1,429
1,425–1,433
–
119 Grok 4.7 (extra-high reasoning)xAI
1,428
1,421–1,435
$2 / $6
119 Grok 4.3xAI
1,428
1,425–1,431
$1.25 / $2.50
119 Qwen3.5-122B-A10BAlibaba
1,428
1,424–1,431
$0.26 / $2.08
122 Inkling SmallThinking Machines
1,426
1,422–1,430
$0.45 / $1.20
123 Gemini 2.5 Flash Preview (09-2025)Google
1,425
1,420–1,429
–
124 Nova Experimental Chat (12-10)Amazon
1,424
1,415–1,434
–
125 Kimi K2 Thinking TurboMoonshot AI
1,423
1,420–1,426
–
126 MiMo-V2-Flash (no reasoning)Xiaomi
1,422
1,419–1,425
–
126 GPT-5 ChatOpenAI
1,422
1,415–1,429
–
126 MiniMax-M2.7MiniMax
1,422
1,419–1,425
$0.30 / $1.20
129 GPT-5 (high reasoning)OpenAI
1,421
1,414–1,428
$1.25 / $10
130 Qwen3.5-27BAlibaba
1,420
1,416–1,424
$0.27 / $2.16
131 Dola-Seed-2.0-ProByteDance
1,419
1,416–1,422
–
131 Qwen3-Next-80B-A3B-InstructAlibaba
1,419
1,412–1,425
$0.10 / $1.10
131 GPT-5.3 ChatOpenAI
1,419
1,415–1,422
–
134 Grok 4 (0709)xAI
1,418
1,412–1,424
–
134 GPT-5.4 nano (high reasoning)OpenAI
1,418
1,415–1,421
$0.20 / $1.25
134 Claude Haiku 4.5Anthropic
1,418
1,416–1,420
$1 / $5
137 Grok 4 Fast (reasoning)xAI
1,417
1,411–1,423
–
137 Step 3.5 FlashStepFun
1,417
1,414–1,420
$0.10 / $0.30
139 Nova Experimental Chat (11-10)Amazon
1,415
1,411–1,419
–
139 Qwen3-VL-235B-A22B-InstructAlibaba
1,415
1,397–1,432
$0.30 / $1.50
141 Qwen3.5-FlashAlibaba
1,414
1,411–1,417
$0.065 / $0.26
142 Grok 4.1 Fast (reasoning)xAI
1,413
1,410–1,416
–
143 Hy3 PreviewTencent
1,412
1,405–1,419
$0.18 / $0.60
144 MiniMax-M2.1 PreviewMiniMax
1,411
1,406–1,416
–
145 MiMo-V2-Flash (reasoning on)Xiaomi
1,410
1,405–1,416
–
145 Solar Pro 4Upstage
1,410
1,405–1,415
$0.09 / $0.36
147 GPT-4.1OpenAI
1,408
1,402–1,415
$2 / $8
148 Qwen3.5-35B-A3BAlibaba
1,407
1,404–1,411
$0.16 / $1.30
148 Gemini 2.5 Flash-Lite Preview (09-2025, no reasoning)Google
1,407
1,404–1,411
–
150 o3OpenAI
1,405
1,398–1,411
$2 / $8
151 Kimi K2 (0905)Moonshot AI
1,403
1,389–1,418
$0.60 / $2.50
151 Nova Experimental Chat (2026-01-10)Amazon
1,403
1,394–1,411
–
153 Muse GlimmerMeta
1,402
1,394–1,410
–
154 Qwen3-Coder-480B-A35BAlibaba
1,401
1,383–1,419
$0.35 / $1.50
155 Nova Experimental Chat (10-20)Amazon
1,399
1,393–1,405
–
156 GLM-4.5-AirZ.ai
1,395
1,388–1,402
$0.14 / $0.86
157 GPT-5 mini (high reasoning)OpenAI
1,394
1,387–1,401
$0.25 / $2
158 GLM-4.6VZ.ai
1,392
1,381–1,402
$0.30 / $0.90
159 Nemotron 3 Super 120B A12BNVIDIA
1,390
1,383–1,397
$0.085 / $0.40
160 MiniMax-M2.5MiniMax
1,386
1,382–1,389
$0.30 / $1.20
161 gpt-oss-120bOpenAI
1,385
1,378–1,391
$0.15 / $0.60
162 Nova 2 LiteAmazon
1,380
1,374–1,386
$0.30 / $2.50
162 Gemma 3 27BGoogle
1,380
1,369–1,391
$0.12 / $0.20
164 Granite 4.2 30BIBM
1,375
1,366–1,384
–
165 INTELLECT-3Prime Intellect
1,373
1,365–1,381
–
165 Trinity Large PreviewArcee AI
1,373
1,369–1,376
–
167 Trinity Large ThinkingArcee AI
1,371
1,367–1,375
$0.25 / $0.80
168 GLM-4.7-FlashZ.ai
1,370
1,364–1,376
$0.06 / $0.40
168 MiniMax-M1MiniMax
1,370
1,359–1,381
$0.40 / $2.20
170 Mercury 2Inception
1,369
1,359–1,378
$0.25 / $0.75
171 o4-miniOpenAI
1,368
1,358–1,378
$1.10 / $4.40
172 Command ACohere
1,363
1,357–1,370
$2.50 / $10
173 Nemotron 3.5 Lightning 30B A3B (NVFP4)NVIDIA
1,362
1,357–1,368
–
173 Nemotron 3 Nano 30B A3BNVIDIA
1,362
1,357–1,367
$0.05 / $0.20
175 MiniMax-M2MiniMax
1,360
1,353–1,367
$0.30 / $1.20
176 Olmo 3.1 32B InstructAi2
1,342
1,336–1,348
–
176 Granite 4.2 8BIBM
1,342
1,332–1,352
$0.06 / $0.25
178 Granite 4.2 3BIBM
1,325
1,315–1,335
–
179 Olmo 3 32B ThinkAi2
1,319
1,312–1,327
–
180 Granite 4.1 8BIBM
1,311
1,302–1,320
–
181 MercuryInception
1,305
1,292–1,317
–
182 Olmo 3.1 32B ThinkAi2
1,300
1,293–1,307
–

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. ≈ marks results whose 95% range overlaps the leader's: they cannot be told apart from it. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Human preference combined with labels of whether each answer's checkable claims were true (factuality weighted 25% by default).

What it does not measure

Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.