Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Estimated effect on users confirming the task was done.

Results dated
28 Sep 2026
Models
46
Unit
IPS effect estimate
Licence
Creative Commons Attribution 4.0 International
Arena Agent: confirmed task success, IPS effect estimate, higher is better
#ModelArena Agent: confirmed task success · 95% range
IPS effect estimate, higher is better
1 Claude Fable 5.1Anthropic
0.17
0.15–0.20
1 Claude Opus 5.5Anthropic
0.16
0.12–0.20
1 GPT-6 AstraOpenAI
0.13
0.09–0.16
1 GPT-6 SolOpenAI
0.11
0.07–0.15
2 DeepSeek V4.1 FlashDeepSeek
0.12
0.11–0.13
2 Claude Opus 5Anthropic
0.10
0.07–0.13
3 Gemini 3.8 FlashGoogle
0.09
0.08–0.11
3 Grok 4.7xAI
0.09
0.06–0.12
4 Kimi K3Moonshot AI
0.09
0.08–0.10
4 Muse Spark 1.3Meta
0.08
0.07–0.10
4 Hy4 previewTencent
0.08
0.06–0.09
4 Claude Opus 5Anthropic
0.07
0.04–0.09
5 GPT-6 LunaOpenAI
0.05
0.01–0.08
6 GLM 5.3 FlashZ.ai
0.07
0.06–0.08
7 GLM 5.3Z.ai
0.06
0.05–0.07
9 Qwen3.8 Flash NextAlibaba
0.06
0.04–0.07
10 Qwen3.8 Max (0902)Alibaba
0.05
0.04–0.07
10 Claude Opus 4.8Anthropic
0.04
0.01–0.06
12 Claude Fable 5Anthropic
0.03
0.01–0.06
12 GPT-5.6 SolOpenAI
0.03
0.01–0.06
13 GLM 5.2Z.ai
0.04
0.02–0.05
15 Claude Sonnet 5Anthropic
0.01
-0.02–0.04
16 Qwen3.8 27BAlibaba
0.02
0.01–0.04
17 DeepSeek V4 Pro 0813DeepSeek
-0.01
-0.05–0.03
18 Gemini 3.7 FlashGoogle
0.00
-0.01–0.02
19 Grok 4.5xAI
-0.01
-0.03–0.01
23 Muse Spark 1.1Meta
-0.02
-0.03–-0.01
23 GPT-5.5OpenAI
-0.03
-0.05–-0.01
23 Grok 4.6xAI
-0.03
-0.06–-0.01
24 GPT-5.4OpenAI
-0.04
-0.05–-0.02
24 DeepSeek V4 Pro 0423DeepSeek
-0.04
-0.06–-0.01
25 GPT-5.5OpenAI
-0.04
-0.06–-0.03
25 GPT-5.6 LunaOpenAI
-0.05
-0.06–-0.03
27 Muse Spark 1.2Meta
-0.05
-0.07–-0.04
27 GPT-5.6 TerraOpenAI
-0.06
-0.09–-0.04
28 Gemini 3.6 FlashGoogle
-0.08
-0.10–-0.05
32 Gemini 3.1 Pro PreviewGoogle
-0.08
-0.10–-0.06
34 Qwen3.7 MaxAlibaba
-0.09
-0.11–-0.07
34 Qwen3.7 PlusAlibaba
-0.09
-0.13–-0.06
35 Hy3Tencent
-0.11
-0.14–-0.09
36 MiMo-V2.5-ProXiaomi
-0.11
-0.13–-0.09
39 MiniMax M3MiniMax
-0.14
-0.17–-0.12
40 Mistral Medium 3.5Mistral
-0.16
-0.20–-0.13
43 Inkling SmallThinkingmachines
-0.21
-0.25–-0.17
43 InklingThinkingmachines
-0.22
-0.25–-0.19
43 Solar Pro 4Upstage
-0.23
-0.28–-0.18

Models share a rank when their ranges overlap. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Estimated effect on users confirming the task was done.

What it does not measure

Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.