Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Estimated effect on the agent taking a user's correction.

Last updated 2 Oct 2026

Results dated
2 Oct 2026
Results
50 configurations of 48 models
Unit
IPS effect estimate
Licence
Creative Commons Attribution 4.0 International

Steerability: Muse Spark 1.3

Top 15 of 50 results · IPS effect estimate, higher is better · lines show the 95% range · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ Gemini 4 Argon (high reasoning)Google 0.13
  2. 2≈ GPT-6 Sol (max reasoning)OpenAI 0.11
  3. 2≈ GPT-6.1 Sol (max reasoning)OpenAI 0.11
  4. 4≈ Claude Opus 5.5 (high reasoning)Anthropic 0.10
  5. 5 Claude Fable 5 (high reasoning)Anthropic 0.09
  6. 5 Claude Fable 5.1 (max reasoning)Anthropic 0.09
  7. 7 GPT-6 Astra (max reasoning)OpenAI 0.08
  8. 8 Claude Sonnet 5.5 (max reasoning)Anthropic 0.07
  9. 8 Claude Opus 5 (high reasoning)Anthropic 0.07
  10. 8 Claude Opus 4.8 (high reasoning)Anthropic 0.07
  11. 11 GPT-5.6 Sol (extra-high reasoning)OpenAI 0.06
  12. 12 Muse Spark 1.3 (max reasoning)Meta 0.05
  13. 12 GPT-5.4 (high reasoning)OpenAI 0.05
  14. 12 Claude Sonnet 5 (high reasoning)Anthropic 0.05
  15. 15 Claude Opus 5 (max reasoning)Anthropic 0.04

Full results

Arena Agent: steerability, IPS effect estimate, higher is better
#ModelSteerability · 95% range
IPS effect estimate, higher is better
Price
$ per million tokens, in / out
1≈ Gemini 4 Argon (high reasoning)Google
0.13
0.12–0.15
–
2≈ GPT-6 Sol (max reasoning)OpenAI
0.11
0.09–0.13
$2 / $10
2≈ GPT-6.1 Sol (max reasoning)OpenAI
0.11
0.08–0.14
$2 / $10
4≈ Claude Opus 5.5 (high reasoning)Anthropic
0.10
0.08–0.13
$4 / $20
5 Claude Fable 5 (high reasoning)Anthropic
0.09
0.08–0.10
$10 / $50
5 Claude Fable 5.1 (max reasoning)Anthropic
0.09
0.06–0.12
$10 / $50
7 GPT-6 Astra (max reasoning)OpenAI
0.08
0.05–0.10
$10 / $50
8 Claude Sonnet 5.5 (max reasoning)Anthropic
0.07
0.04–0.10
$2 / $10
8 Claude Opus 5 (high reasoning)Anthropic
0.07
0.05–0.08
$5 / $25
8 Claude Opus 4.8 (high reasoning)Anthropic
0.07
0.05–0.08
$5 / $25
11 GPT-5.6 Sol (extra-high reasoning)OpenAI
0.06
0.04–0.08
$4 / $20
12 Muse Spark 1.3 (max reasoning)Meta
0.05
0.05–0.06
$1.25 / $4.25
12 GPT-5.4 (high reasoning)OpenAI
0.05
0.04–0.06
$2.50 / $15
12 Claude Sonnet 5 (high reasoning)Anthropic
0.05
0.03–0.07
$2 / $10
15 Claude Opus 5 (max reasoning)Anthropic
0.04
0.02–0.06
$5 / $25
15 GPT-5.5 (extra-high reasoning)OpenAI
0.04
0.03–0.06
$5 / $30
15 GPT-6 Luna (max reasoning)OpenAI
0.04
0.03–0.05
$0.10 / $0.50
18 Gemini 3.8 Flash (high reasoning)Google
0.03
0.02–0.04
$1.50 / $7.50
18 Qwen3.8-Max (0902)Alibaba
0.03
0.02–0.04
$2 / $6
20 GPT-5.6 Terra (extra-high reasoning)OpenAI
0.02
0.01–0.04
$2 / $12
20 Step 5 PreviewStepFun
0.02
-0.00–0.04
–
20 Grok 4.5xAI
0.02
0.01–0.03
$2 / $6
23 GLM-5.2 (max reasoning)Z.ai
0.01
0.01–0.02
$1.40 / $4.40
23 Kimi K3 (max reasoning)Moonshot AI
0.01
0.01–0.02
$3 / $15
23 GPT-5.5OpenAI
0.01
0.00–0.03
$5 / $30
23 DeepSeek-V4.1-Flash (max reasoning)DeepSeek
0.01
0.01–0.02
$0.30 / $1.20
23 DeepSeek-V4-Pro (0813, high reasoning)DeepSeek
0.01
-0.01–0.03
$1.32 / $3.96
23 GLM-5.3 (max reasoning)Z.ai
0.01
-0.00–0.02
$1.40 / $4.40
23 Hy4 PreviewTencent
0.01
-0.00–0.02
$0.83 / $2.50
23 Hy3Tencent
0.01
-0.01–0.02
$0.14 / $0.58
23 Grok 4.6 (extra-high reasoning)xAI
0.01
-0.01–0.02
$2 / $6
32 Grok 4.7 (extra-high reasoning)xAI
0.00
-0.02–0.03
$2 / $6
32 GPT-5.6 Luna (extra-high reasoning)OpenAI
0.00
-0.01–0.01
$0.20 / $1.20
32 MiMo-V2.6-FlashXiaomi
-0.00
-0.02–0.01
$0.14 / $0.28
32 Qwen3.8-27BAlibaba
-0.00
-0.01–0.00
$0.50 / $3
36 GLM-5.3-FlashZ.ai
-0.01
-0.01–0.00
$0.15 / $0.50
36 Qwen3.8-Flash-NextAlibaba
-0.01
-0.02–0.00
–
36 Gemini 3.7 Flash (high reasoning)Google
-0.01
-0.02–-0.00
$1.50 / $7.50
39 Gemini 3.1 Pro PreviewGoogle
-0.05
-0.06–-0.03
$2 / $12
39 Muse Spark 1.2 (extra-high reasoning)Meta
-0.05
-0.06–-0.04
$1.25 / $4.25
41 Mistral Medium 3.5Mistral AI
-0.06
-0.09–-0.04
$1.50 / $7.50
41 Muse Spark 1.1Meta
-0.06
-0.07–-0.06
$1.25 / $4.25
41 Qwen3.7-MaxAlibaba
-0.06
-0.08–-0.05
$1.48 / $4.42
44 MiniMax-M3MiniMax
-0.07
-0.09–-0.06
$0.30 / $1.20
44 Gemini 3.6 Flash (high reasoning)Google
-0.07
-0.09–-0.06
$1.50 / $7.50
46 MiMo-V2.5-ProXiaomi
-0.08
-0.09–-0.06
$0.43 / $0.87
46 Qwen3.7-PlusAlibaba
-0.08
-0.10–-0.06
$0.32 / $1.28
48 Solar Pro 4Upstage
-0.09
-0.12–-0.06
$0.09 / $0.36
49 InklingThinking Machines
-0.12
-0.14–-0.10
$0.95 / $4.05
50 Inkling SmallThinking Machines
-0.14
-0.17–-0.11
$0.45 / $1.20

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. ≈ marks results whose 95% range overlaps the leader's: they cannot be told apart from it. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Estimated effect on the agent taking a user's correction.

What it does not measure

Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.