Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Estimated effect on users confirming the task was done.

Last updated 2 Oct 2026

Results dated
2 Oct 2026
Results
50 configurations of 48 models
Unit
IPS effect estimate
Licence
Creative Commons Attribution 4.0 International

Confirmed task success: Gemini 3.8 Flash

Top 15 of 48 results · IPS effect estimate, higher is better · lines show the 95% range · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ Claude Fable 5.1 (max reasoning)Anthropic 0.18
  2. 2≈ GPT-6.1 Sol (max reasoning)OpenAI 0.15
  3. 2≈ Gemini 4 Argon (high reasoning)Google 0.15
  4. 4≈ Claude Opus 5.5 (high reasoning)Anthropic 0.14
  5. 4≈ Claude Sonnet 5.5 (max reasoning)Anthropic 0.14
  6. 6≈ GPT-6 Astra (max reasoning)OpenAI 0.13
  7. 7 Claude Opus 5 (max reasoning)Anthropic 0.10
  8. 8 DeepSeek-V4.1-Flash (max reasoning)DeepSeek 0.08
  9. 8 Kimi K3 (max reasoning)Moonshot AI 0.08
  10. 10 GPT-6 Sol (max reasoning)OpenAI 0.07
  11. 10 Gemini 3.8 Flash (high reasoning)Google 0.07
  12. 12 Claude Fable 5 (high reasoning)Anthropic 0.06
  13. 12 Hy4 PreviewTencent 0.06
  14. 12 MiMo-V2.6-FlashXiaomi 0.06
  15. 15 Grok 4.7 (extra-high reasoning)xAI 0.05

Full results

Arena Agent: confirmed task success, IPS effect estimate, higher is better
#ModelConfirmed task success · 95% range
IPS effect estimate, higher is better
Price
$ per million tokens, in / out
1≈ Claude Fable 5.1 (max reasoning)Anthropic
0.18
0.15–0.20
$10 / $50
2≈ GPT-6.1 Sol (max reasoning)OpenAI
0.15
0.11–0.20
$2 / $10
2≈ Gemini 4 Argon (high reasoning)Google
0.15
0.12–0.18
–
4≈ Claude Opus 5.5 (high reasoning)Anthropic
0.14
0.10–0.18
$4 / $20
4≈ Claude Sonnet 5.5 (max reasoning)Anthropic
0.14
0.09–0.19
$2 / $10
6≈ GPT-6 Astra (max reasoning)OpenAI
0.13
0.10–0.17
$10 / $50
7 Claude Opus 5 (max reasoning)Anthropic · best of 2 settings
0.10
0.06–0.13
$5 / $25
8 DeepSeek-V4.1-Flash (max reasoning)DeepSeek
0.08
0.07–0.09
$0.30 / $1.20
8 Kimi K3 (max reasoning)Moonshot AI
0.08
0.06–0.09
$3 / $15
10 GPT-6 Sol (max reasoning)OpenAI
0.07
0.03–0.11
$2 / $10
10 Gemini 3.8 Flash (high reasoning)Google
0.07
0.05–0.09
$1.50 / $7.50
12 Claude Fable 5 (high reasoning)Anthropic
0.06
0.03–0.09
$10 / $50
12 Hy4 PreviewTencent
0.06
0.04–0.07
$0.83 / $2.50
12 MiMo-V2.6-FlashXiaomi
0.06
0.03–0.08
$0.14 / $0.28
15 Grok 4.7 (extra-high reasoning)xAI
0.05
0.02–0.09
$2 / $6
15 GLM-5.3 (max reasoning)Z.ai
0.05
0.04–0.07
$1.40 / $4.40
15 GPT-5.6 Sol (extra-high reasoning)OpenAI
0.05
0.02–0.08
$4 / $20
15 Muse Spark 1.3 (max reasoning)Meta
0.05
0.04–0.06
$1.25 / $4.25
15 GLM-5.2 (max reasoning)Z.ai
0.05
0.03–0.06
$1.40 / $4.40
15 Qwen3.8-Max (0902)Alibaba
0.05
0.03–0.06
$2 / $6
15 Step 5 PreviewStepFun
0.05
0.01–0.08
–
22 GLM-5.3-FlashZ.ai
0.04
0.03–0.05
$0.15 / $0.50
22 Claude Opus 4.8 (high reasoning)Anthropic
0.04
0.02–0.07
$5 / $25
22 Qwen3.8-Flash-NextAlibaba
0.04
0.02–0.05
–
25 Grok 4.5xAI
0.01
-0.01–0.04
$2 / $6
25 Claude Sonnet 5 (high reasoning)Anthropic
0.01
-0.02–0.05
$2 / $10
25 Gemini 3.7 Flash (high reasoning)Google
0.01
-0.01–0.03
$1.50 / $7.50
25 Qwen3.8-27BAlibaba
0.01
-0.00–0.02
$0.50 / $3
29 GPT-6 Luna (max reasoning)OpenAI
-0.00
-0.03–0.02
$0.10 / $0.50
30 Grok 4.6 (extra-high reasoning)xAI
-0.01
-0.03–0.02
$2 / $6
30 GPT-5.5 (extra-high reasoning)OpenAI · best of 2 settings
-0.01
-0.03–0.01
$5 / $30
32 DeepSeek-V4-Pro (0813, high reasoning)DeepSeek
-0.02
-0.06–0.02
$1.32 / $3.96
33 GPT-5.4 (high reasoning)OpenAI
-0.03
-0.05–-0.00
$2.50 / $15
33 GPT-5.6 Luna (extra-high reasoning)OpenAI
-0.03
-0.05–-0.02
$0.20 / $1.20
35 GPT-5.6 Terra (extra-high reasoning)OpenAI
-0.04
-0.06–-0.01
$2 / $12
35 Muse Spark 1.1Meta
-0.04
-0.05–-0.03
$1.25 / $4.25
37 Muse Spark 1.2 (extra-high reasoning)Meta
-0.05
-0.07–-0.04
$1.25 / $4.25
38 Gemini 3.6 Flash (high reasoning)Google
-0.06
-0.08–-0.03
$1.50 / $7.50
39 Qwen3.7-MaxAlibaba
-0.09
-0.12–-0.07
$1.48 / $4.42
39 Qwen3.7-PlusAlibaba
-0.09
-0.13–-0.06
$0.32 / $1.28
41 MiMo-V2.5-ProXiaomi
-0.10
-0.12–-0.07
$0.43 / $0.87
41 Gemini 3.1 Pro PreviewGoogle
-0.10
-0.12–-0.08
$2 / $12
43 Hy3Tencent
-0.11
-0.14–-0.08
$0.14 / $0.58
44 MiniMax-M3MiniMax
-0.14
-0.16–-0.11
$0.30 / $1.20
45 Mistral Medium 3.5Mistral AI
-0.15
-0.18–-0.11
$1.50 / $7.50
46 Inkling SmallThinking Machines
-0.20
-0.24–-0.16
$0.45 / $1.20
47 Solar Pro 4Upstage
-0.21
-0.26–-0.15
$0.09 / $0.36
47 InklingThinking Machines
-0.21
-0.24–-0.18
$0.95 / $4.05

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. ≈ marks results whose 95% range overlaps the leader's: they cannot be told apart from it. Each model is shown at its best setting; show every setting. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Estimated effect on users confirming the task was done.

What it does not measure

Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.