Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Estimated effect on the agent taking a user's correction.

Results dated
28 Sep 2026
Models
46
Unit
IPS effect estimate
Licence
Creative Commons Attribution 4.0 International
Arena Agent: steerability, IPS effect estimate, higher is better
#ModelArena Agent: steerability · 95% range
IPS effect estimate, higher is better
1 GPT-6 SolOpenAI
0.15
0.10–0.19
1 Claude Opus 5.5Anthropic
0.14
0.10–0.18
1 Claude Fable 5Anthropic
0.11
0.09–0.14
1 Claude Opus 5Anthropic
0.11
0.08–0.13
1 Claude Opus 4.8Anthropic
0.10
0.07–0.12
1 Claude Fable 5.1Anthropic
0.08
0.05–0.12
1 Claude Opus 5Anthropic
0.08
0.05–0.11
1 Claude Sonnet 5Anthropic
0.07
0.04–0.11
4 GPT-5.5OpenAI
0.07
0.05–0.09
4 GPT-5.6 SolOpenAI
0.07
0.04–0.09
6 Gemini 3.8 FlashGoogle
0.04
0.02–0.06
6 GLM 5.2Z.ai
0.04
0.03–0.06
6 GPT-5.6 TerraOpenAI
0.04
0.02–0.06
6 Grok 4.5xAI
0.04
0.02–0.06
6 Grok 4.6xAI
0.03
0.01–0.05
6 DeepSeek V4 Pro 0813DeepSeek
0.02
-0.02–0.06
6 GPT-6 LunaOpenAI
0.02
-0.02–0.06
10 Grok 4.7xAI
0.00
-0.03–0.04
11 GLM 5.3Z.ai
0.03
0.02–0.04
11 GPT-5.4OpenAI
0.02
0.01–0.04
11 GPT-5.5OpenAI
0.02
0.01–0.04
11 Qwen3.8 Max (0902)Alibaba
0.02
0.01–0.03
11 GPT-5.6 LunaOpenAI
0.02
0.01–0.03
11 Muse Spark 1.3Meta
0.02
0.00–0.03
11 GPT-6 AstraOpenAI
-0.01
-0.06–0.04
16 Gemini 3.7 FlashGoogle
-0.00
-0.01–0.01
17 Hy4 previewTencent
-0.01
-0.03–0.01
17 Hy3Tencent
-0.01
-0.04–0.01
20 Kimi K3Moonshot AI
-0.01
-0.02–0.01
21 Qwen3.8 27BAlibaba
-0.01
-0.02–0.00
22 GLM 5.3 FlashZ.ai
-0.01
-0.02–-0.00
22 DeepSeek V4 Pro 0423DeepSeek
-0.03
-0.05–-0.01
24 DeepSeek V4.1 FlashDeepSeek
-0.03
-0.05–-0.02
24 Qwen3.8 Flash NextAlibaba
-0.03
-0.05–-0.02
24 Qwen3.7 MaxAlibaba
-0.03
-0.05–-0.02
24 Gemini 3.6 FlashGoogle
-0.04
-0.06–-0.02
29 Gemini 3.1 Pro PreviewGoogle
-0.04
-0.05–-0.03
29 MiMo-V2.5-ProXiaomi
-0.05
-0.06–-0.03
31 Muse Spark 1.2Meta
-0.05
-0.07–-0.04
31 Mistral Medium 3.5Mistral
-0.07
-0.10–-0.04
33 Muse Spark 1.1Meta
-0.06
-0.07–-0.05
34 Qwen3.7 PlusAlibaba
-0.07
-0.10–-0.05
35 MiniMax M3MiniMax
-0.07
-0.09–-0.05
40 Solar Pro 4Upstage
-0.10
-0.14–-0.07
41 Inkling SmallThinkingmachines
-0.12
-0.15–-0.08
42 InklingThinkingmachines
-0.12
-0.14–-0.10

Models share a rank when their ranges overlap. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Estimated effect on the agent taking a user's correction.

What it does not measure

Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.