Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Estimated effect on users praising rather than complaining.

Last updated 2 Oct 2026

Results dated
2 Oct 2026
Results
50 configurations of 48 models
Unit
IPS effect estimate
Licence
Creative Commons Attribution 4.0 International

Praise over complaint: Claude Sonnet 5.5

Top 15 of 48 results · IPS effect estimate, higher is better · lines show the 95% range · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ GPT-6 Astra (max reasoning)OpenAI 0.35
  2. 2≈ Claude Fable 5.1 (max reasoning)Anthropic 0.32
  3. 3≈ Claude Opus 5.5 (high reasoning)Anthropic 0.31
  4. 4≈ Gemini 4 Argon (high reasoning)Google 0.28
  5. 5≈ Claude Sonnet 5.5 (max reasoning)Anthropic 0.26
  6. 5≈ GPT-6.1 Sol (max reasoning)OpenAI 0.26
  7. 7≈ GPT-6 Sol (max reasoning)OpenAI 0.23
  8. 8 GPT-5.6 Sol (extra-high reasoning)OpenAI 0.20
  9. 9 Claude Fable 5 (high reasoning)Anthropic 0.17
  10. 9 Claude Opus 5 (high reasoning)Anthropic 0.17
  11. 11 Claude Opus 4.8 (high reasoning)Anthropic 0.15
  12. 12 GLM-5.2 (max reasoning)Z.ai 0.10
  13. 13 Claude Sonnet 5 (high reasoning)Anthropic 0.09
  14. 14 GPT-5.5 (extra-high reasoning)OpenAI 0.08
  15. 15 Grok 4.7 (extra-high reasoning)xAI 0.07

Full results

Arena Agent: praise over complaint, IPS effect estimate, higher is better
#ModelPraise over complaint · 95% range
IPS effect estimate, higher is better
Price
$ per million tokens, in / out
1≈ GPT-6 Astra (max reasoning)OpenAI
0.35
0.25–0.44
$10 / $50
2≈ Claude Fable 5.1 (max reasoning)Anthropic
0.32
0.24–0.40
$10 / $50
3≈ Claude Opus 5.5 (high reasoning)Anthropic
0.31
0.23–0.40
$4 / $20
4≈ Gemini 4 Argon (high reasoning)Google
0.28
0.20–0.35
–
5≈ Claude Sonnet 5.5 (max reasoning)Anthropic
0.26
0.13–0.39
$2 / $10
5≈ GPT-6.1 Sol (max reasoning)OpenAI
0.26
0.13–0.38
$2 / $10
7≈ GPT-6 Sol (max reasoning)OpenAI
0.23
0.13–0.33
$2 / $10
8 GPT-5.6 Sol (extra-high reasoning)OpenAI
0.20
0.15–0.24
$4 / $20
9 Claude Fable 5 (high reasoning)Anthropic
0.17
0.13–0.22
$10 / $50
9 Claude Opus 5 (high reasoning)Anthropic · best of 2 settings
0.17
0.11–0.22
$5 / $25
11 Claude Opus 4.8 (high reasoning)Anthropic
0.15
0.10–0.20
$5 / $25
12 GLM-5.2 (max reasoning)Z.ai
0.10
0.07–0.13
$1.40 / $4.40
13 Claude Sonnet 5 (high reasoning)Anthropic
0.09
0.03–0.14
$2 / $10
14 GPT-5.5 (extra-high reasoning)OpenAI · best of 2 settings
0.08
0.04–0.11
$5 / $30
15 Grok 4.7 (extra-high reasoning)xAI
0.07
0.01–0.13
$2 / $6
15 Kimi K3 (max reasoning)Moonshot AI
0.07
0.05–0.09
$3 / $15
15 Hy4 PreviewTencent
0.07
0.04–0.10
$0.83 / $2.50
18 GLM-5.3 (max reasoning)Z.ai
0.05
0.03–0.08
$1.40 / $4.40
18 Gemini 3.8 Flash (high reasoning)Google
0.05
0.02–0.08
$1.50 / $7.50
20 Qwen3.8-Max (0902)Alibaba
0.03
0.01–0.05
$2 / $6
20 Muse Spark 1.3 (max reasoning)Meta
0.03
0.01–0.05
$1.25 / $4.25
20 GPT-5.6 Terra (extra-high reasoning)OpenAI
0.03
-0.01–0.07
$2 / $12
23 DeepSeek-V4.1-Flash (max reasoning)DeepSeek
0.02
0.01–0.04
$0.30 / $1.20
23 DeepSeek-V4-Pro (0813, high reasoning)DeepSeek
0.02
-0.04–0.08
$1.32 / $3.96
25 GPT-5.4 (high reasoning)OpenAI
0.00
-0.03–0.04
$2.50 / $15
25 Grok 4.6 (extra-high reasoning)xAI
0.00
-0.03–0.04
$2 / $6
27 MiMo-V2.6-FlashXiaomi
-0.01
-0.05–0.03
$0.14 / $0.28
27 GLM-5.3-FlashZ.ai
-0.01
-0.02–0.01
$0.15 / $0.50
27 Gemini 3.1 Pro PreviewGoogle
-0.01
-0.04–0.01
$2 / $12
30 GPT-6 Luna (max reasoning)OpenAI
-0.02
-0.05–0.02
$0.10 / $0.50
31 Qwen3.8-27BAlibaba
-0.03
-0.05–-0.00
$0.50 / $3
31 Grok 4.5xAI
-0.03
-0.06–0.00
$2 / $6
31 Gemini 3.7 Flash (high reasoning)Google
-0.03
-0.06–-0.01
$1.50 / $7.50
34 Qwen3.8-Flash-NextAlibaba
-0.04
-0.06–-0.01
–
34 GPT-5.6 Luna (extra-high reasoning)OpenAI
-0.04
-0.06–-0.02
$0.20 / $1.20
36 Step 5 PreviewStepFun
-0.05
-0.10–0.00
–
37 Hy3Tencent
-0.07
-0.10–-0.03
$0.14 / $0.58
38 Gemini 3.6 Flash (high reasoning)Google
-0.09
-0.12–-0.05
$1.50 / $7.50
38 Qwen3.7-MaxAlibaba
-0.09
-0.12–-0.07
$1.48 / $4.42
40 Qwen3.7-PlusAlibaba
-0.11
-0.15–-0.07
$0.32 / $1.28
41 Muse Spark 1.1Meta
-0.12
-0.13–-0.11
$1.25 / $4.25
41 Muse Spark 1.2 (extra-high reasoning)Meta
-0.12
-0.14–-0.10
$1.25 / $4.25
41 MiMo-V2.5-ProXiaomi
-0.12
-0.15–-0.10
$0.43 / $0.87
44 MiniMax-M3MiniMax
-0.14
-0.17–-0.12
$0.30 / $1.20
45 Mistral Medium 3.5Mistral AI
-0.18
-0.21–-0.14
$1.50 / $7.50
46 Inkling SmallThinking Machines
-0.22
-0.26–-0.19
$0.45 / $1.20
47 InklingThinking Machines
-0.23
-0.25–-0.20
$0.95 / $4.05
47 Solar Pro 4Upstage
-0.23
-0.28–-0.19
$0.09 / $0.36

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. ≈ marks results whose 95% range overlaps the leader's: they cannot be told apart from it. Each model is shown at its best setting; show every setting. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Estimated effect on users praising rather than complaining.

What it does not measure

Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.