Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Estimated effect on avoiding calls to tools that do not exist.

Last updated 2 Oct 2026

Results dated
2 Oct 2026
Results
50 configurations of 48 models
Unit
IPS effect estimate
Licence
Creative Commons Attribution 4.0 International

Tool grounding: GLM-5.3

Top 15 of 48 results · IPS effect estimate, higher is better · lines show the 95% range · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ GPT-6.1 Sol (max reasoning)OpenAI 0.00
  2. 1≈ Kimi K3 (max reasoning)Moonshot AI 0.00
  3. 1≈ Claude Fable 5 (high reasoning)Anthropic 0.00
  4. 1≈ Claude Fable 5.1 (max reasoning)Anthropic 0.00
  5. 1≈ Claude Opus 5 (high reasoning)Anthropic 0.00
  6. 1≈ Claude Opus 5.5 (high reasoning)Anthropic 0.00
  7. 1≈ Claude Sonnet 5.5 (max reasoning)Anthropic 0.00
  8. 1≈ Step 5 PreviewStepFun 0.00
  9. 1≈ GLM-5.2 (max reasoning)Z.ai 0.00
  10. 1≈ Gemini 3.6 Flash (high reasoning)Google 0.00
  11. 1≈ Grok 4.5xAI 0.00
  12. 1≈ GLM-5.3 (max reasoning)Z.ai 0.00
  13. 1≈ GPT-6 Astra (max reasoning)OpenAI 0.00
  14. 1≈ Muse Spark 1.3 (max reasoning)Meta 0.00
  15. 1≈ GLM-5.3-FlashZ.ai 0.00

Full results

Arena Agent: tool grounding, IPS effect estimate, higher is better
#ModelTool grounding · 95% range
IPS effect estimate, higher is better
Price
$ per million tokens, in / out
1≈ GPT-6.1 Sol (max reasoning)OpenAI
0.00
0.00–0.01
$2 / $10
1≈ Kimi K3 (max reasoning)Moonshot AI
0.00
0.00–0.01
$3 / $15
1≈ Claude Fable 5 (high reasoning)Anthropic
0.00
0.00–0.01
$10 / $50
1≈ Claude Fable 5.1 (max reasoning)Anthropic
0.00
0.00–0.01
$10 / $50
1≈ Claude Opus 5 (high reasoning)Anthropic · best of 2 settings
0.00
0.00–0.01
$5 / $25
1≈ Claude Opus 5.5 (high reasoning)Anthropic
0.00
0.00–0.01
$4 / $20
1≈ Claude Sonnet 5.5 (max reasoning)Anthropic
0.00
0.00–0.01
$2 / $10
1≈ Step 5 PreviewStepFun
0.00
0.00–0.01
–
1≈ GLM-5.2 (max reasoning)Z.ai
0.00
0.00–0.00
$1.40 / $4.40
1≈ Gemini 3.6 Flash (high reasoning)Google
0.00
0.00–0.00
$1.50 / $7.50
1≈ Grok 4.5xAI
0.00
0.00–0.01
$2 / $6
1≈ GLM-5.3 (max reasoning)Z.ai
0.00
0.00–0.01
$1.40 / $4.40
1≈ GPT-6 Astra (max reasoning)OpenAI
0.00
0.00–0.01
$10 / $50
1≈ Muse Spark 1.3 (max reasoning)Meta
0.00
0.00–0.00
$1.25 / $4.25
1≈ GLM-5.3-FlashZ.ai
0.00
0.00–0.00
$0.15 / $0.50
1≈ DeepSeek-V4.1-Flash (max reasoning)DeepSeek
0.00
0.00–0.00
$0.30 / $1.20
1≈ Gemini 3.8 Flash (high reasoning)Google
0.00
0.00–0.00
$1.50 / $7.50
1≈ GPT-5.5 (extra-high reasoning)OpenAI · best of 2 settings
0.00
0.00–0.00
$5 / $30
1≈ Claude Sonnet 5 (high reasoning)Anthropic
0.00
0.00–0.00
$2 / $10
1≈ Grok 4.6 (extra-high reasoning)xAI
0.00
0.00–0.00
$2 / $6
1≈ DeepSeek-V4-Pro (0813, high reasoning)DeepSeek
0.00
0.00–0.00
$1.32 / $3.96
1≈ GPT-5.6 Luna (extra-high reasoning)OpenAI
0.00
0.00–0.00
$0.20 / $1.20
1≈ Grok 4.7 (extra-high reasoning)xAI
0.00
0.00–0.00
$2 / $6
1≈ Claude Opus 4.8 (high reasoning)Anthropic
0.00
0.00–0.00
$5 / $25
1≈ Gemini 3.7 Flash (high reasoning)Google
0.00
0.00–0.00
$1.50 / $7.50
1≈ GPT-5.6 Sol (extra-high reasoning)OpenAI
0.00
0.00–0.00
$4 / $20
1≈ InklingThinking Machines
0.00
0.00–0.00
$0.95 / $4.05
1≈ GPT-5.4 (high reasoning)OpenAI
0.00
-0.00–0.00
$2.50 / $15
1≈ Gemini 4 Argon (high reasoning)Google
0.00
-0.00–0.00
–
1≈ GPT-6 Sol (max reasoning)OpenAI
0.00
-0.00–0.00
$2 / $10
1≈ Hy4 PreviewTencent
0.00
-0.00–0.00
$0.83 / $2.50
1≈ MiniMax-M3MiniMax
0.00
-0.00–0.00
$0.30 / $1.20
1≈ GPT-5.6 Terra (extra-high reasoning)OpenAI
0.00
-0.00–0.00
$2 / $12
1≈ Qwen3.7-MaxAlibaba
0.00
-0.00–0.00
$1.48 / $4.42
1≈ Gemini 3.1 Pro PreviewGoogle
-0.00
-0.01–0.01
$2 / $12
1≈ Muse Spark 1.2 (extra-high reasoning)Meta
-0.00
-0.00–-0.00
$1.25 / $4.25
1≈ Muse Spark 1.1Meta
-0.00
-0.00–-0.00
$1.25 / $4.25
1≈ Qwen3.8-27BAlibaba
-0.00
-0.00–-0.00
$0.50 / $3
1≈ GPT-6 Luna (max reasoning)OpenAI
-0.00
-0.00–-0.00
$0.10 / $0.50
1≈ Inkling SmallThinking Machines
-0.00
-0.01–-0.00
$0.45 / $1.20
41 Qwen3.8-Flash-NextAlibaba
-0.01
-0.01–-0.01
–
41 Qwen3.7-PlusAlibaba
-0.01
-0.01–-0.00
$0.32 / $1.28
41 Qwen3.8-Max (0902)Alibaba
-0.01
-0.01–-0.01
$2 / $6
41 MiMo-V2.5-ProXiaomi
-0.01
-0.01–-0.01
$0.43 / $0.87
41 Solar Pro 4Upstage
-0.01
-0.02–-0.01
$0.09 / $0.36
46 Mistral Medium 3.5Mistral AI
-0.02
-0.03–-0.01
$1.50 / $7.50
47 Hy3Tencent
-0.03
-0.04–-0.02
$0.14 / $0.58
48 MiMo-V2.6-FlashXiaomi
-0.04
-0.04–-0.03
$0.14 / $0.28

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. ≈ marks results whose 95% range overlaps the leader's: they cannot be told apart from it. Each model is shown at its best setting; show every setting. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Estimated effect on avoiding calls to tools that do not exist.

What it does not measure

Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.