Benchmarks / Artificial Analysis

Reported by Artificial Analysis

Artificial Analysis

Share of simulated bank customer-service tasks an agent resolves with tools, as run by Artificial Analysis with its own harness and prompts.

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Results
175 configurations of 121 models
Unit
% of tasks
Licence
Artificial Analysis commercial data licence

τ-bench banking: GLM-5.3

Top 15 of 121 results · % of tasks, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 Qwen3.8-Max (0803)Alibaba 51.3%
  2. 2 Grok 4.6 (high reasoning)xAI 50.7%
  3. 3 Muse Spark 1.3 (max reasoning)Meta 50.5%
  4. 4 GLM-5.3 (max reasoning)Z.ai 50.3%
  5. 5 Qwen3.8-2.4T-A95BAlibaba 49.1%
  6. 6 Qwen3.8-27B (extra-high reasoning)Alibaba 48.0%
  7. 7 Qwen3.8-Max (0902)Alibaba 47.8%
  8. 8 Claude Fable 5.1 (max reasoning)Anthropic 47.2%
  9. 8 GLM-5.3-FlashZ.ai 47.2%
  10. 10 Kimi K3 (max reasoning)Moonshot AI 46.0%
  11. 11 Gemini 3.8 Flash (medium reasoning)Google 45.8%
  12. 12 Qwen3.8-Flash-NextAlibaba 45.4%
  13. 13 Claude Opus 5 (high reasoning)Anthropic 44.7%
  14. 14 GPT-5.6 Sol (max reasoning)OpenAI 44.3%
  15. 15 GPT-6 Astra (extra-high reasoning)OpenAI 43.1%

Full results

Artificial Analysis: τ-bench banking, % of tasks, higher is better
#Modelτ-bench banking
% of tasks, higher is better
Price
$ per million tokens, in / out
1 Qwen3.8-Max (0803)Alibaba
51.3%
–
2 Grok 4.6 (high reasoning)xAI · best of 4 settings
50.7%
$2 / $6
3 Muse Spark 1.3 (max reasoning)Meta · best of 2 settings
50.5%
$1.25 / $4.25
4 GLM-5.3 (max reasoning)Z.ai
50.3%
$1.40 / $4.40
5 Qwen3.8-2.4T-A95BAlibaba
49.1%
$2 / $6
6 Qwen3.8-27B (extra-high reasoning)Alibaba · best of 4 settings
48.0%
$0.50 / $3
7 Qwen3.8-Max (0902)Alibaba
47.8%
$2 / $6
8 Claude Fable 5.1 (max reasoning)Anthropic · best of 5 settings
47.2%
$10 / $50
8 GLM-5.3-FlashZ.ai
47.2%
$0.15 / $0.50
10 Kimi K3 (max reasoning)Moonshot AI · best of 2 settings
46.0%
$3 / $15
11 Gemini 3.8 Flash (medium reasoning)Google · best of 3 settings
45.8%
$1.50 / $7.50
12 Qwen3.8-Flash-NextAlibaba
45.4%
–
13 Claude Opus 5 (high reasoning)Anthropic · best of 5 settings
44.7%
$5 / $25
14 GPT-5.6 Sol (max reasoning)OpenAI · best of 6 settings
44.3%
$4 / $20
15 GPT-6 Astra (extra-high reasoning)OpenAI · best of 5 settings
43.1%
$10 / $50
16 Grok 4.5 (high reasoning)xAI
42.1%
$2 / $6
17 DeepSeek-V4-Flash-Vision-Exp (max reasoning)DeepSeek
41.0%
$0.44 / $1.32
18 GPT-5.6 Terra (max reasoning)OpenAI · best of 6 settings
40.2%
$2 / $12
19 DeepSeek-V4-Pro (0813, max reasoning)DeepSeek
39.6%
$1.32 / $3.96
19 GPT-5.4 (extra-high reasoning)OpenAI
39.6%
$2.50 / $15
21 DeepSeek-V4-Flash (0731, max reasoning)DeepSeek
39.4%
$0.14 / $0.28
22 GPT-5.5 (extra-high reasoning)OpenAI · best of 5 settings
39.0%
$5 / $30
23 Claude Fable 5 (max reasoning)Anthropic
38.1%
$10 / $50
24 Claude Sonnet 5 (max reasoning)Anthropic · best of 2 settings
37.3%
$2 / $10
25 Gemini 3.7 Flash (medium reasoning)Google · best of 3 settings
35.5%
$1.50 / $7.50
26 Muse Spark 1.2 (extra-high reasoning)Meta
34.8%
$1.25 / $4.25
27 Claude Opus 4.7 (max reasoning)Anthropic
34.6%
$5 / $25
27 GLM-5.2 (max reasoning)Z.ai · best of 2 settings
34.6%
$1.40 / $4.40
29 Claude Sonnet 4.6 (max reasoning)Anthropic
34.4%
$3 / $15
30 Claude Opus 4.8 (max reasoning)Anthropic
34.2%
$5 / $25
31 Gemini 3.5 Flash (high reasoning)Google
32.2%
$1.50 / $9
32 Muse Spark 1.1 (extra-high reasoning)Meta
31.8%
$1.25 / $4.25
33 GPT-5.6 Luna (max reasoning)OpenAI · best of 6 settings
31.1%
$0.20 / $1.20
34 DeepSeek-V4-Flash (0423, max reasoning)DeepSeek · best of 2 settings
30.9%
$0.14 / $0.28
35 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek · best of 2 settings
30.1%
$1.42 / $2.83
36 Gemini 3.6 Flash (high reasoning)Google
29.9%
$1.50 / $7.50
37 Inkling (extra-high reasoning)Thinking Machines
29.1%
$0.95 / $4.05
38 GPT-5.4 nano (extra-high reasoning)OpenAI
27.4%
$0.20 / $1.25
39 GPT-5.4 mini (extra-high reasoning)OpenAI
25.6%
$0.75 / $4.50
40 Claude Sonnet 4.5 (reasoning on)Anthropic
24.5%
$3 / $15
41 Muse Glimmer (high reasoning)Meta
23.5%
–
42 Kimi K2.6Moonshot AI
23.3%
$0.95 / $4
42 Solar Pro 4Upstage
23.3%
$0.09 / $0.36
44 Hy3Tencent
22.9%
$0.14 / $0.58
45 GPT-5 (high reasoning)OpenAI
22.1%
$1.25 / $10
46 Gemini 3.1 Pro PreviewGoogle
21.4%
$2 / $12
47 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek
21.0%
$0.27 / $1
48 Qwen3.6-PlusAlibaba
20.8%
$0.33 / $1.95
48 Gemini 3 Flash Preview (reasoning on)Google
20.8%
$0.50 / $3
50 Kimi K2.7 CodeMoonshot AI
20.2%
$0.95 / $4
51 Inkling SmallThinking Machines
18.8%
$0.45 / $1.20
52 Qwen3.7-PlusAlibaba
17.5%
$0.32 / $1.28
52 Gemini 3.5 Flash-LiteGoogle
17.5%
$0.30 / $2.50
54 Qwen3.6-27BAlibaba · best of 2 settings
16.7%
$0.30 / $3.20
54 Claude Sonnet 4 (reasoning on)Anthropic
16.7%
$3 / $15
56 GPT-5.1 (high reasoning)OpenAI
15.9%
$1.25 / $10
57 GPT-5 mini (high reasoning)OpenAI
15.5%
$0.25 / $2
58 Qwen3.5-122B-A10BAlibaba · best of 2 settings
15.3%
$0.26 / $2.08
58 MiniMax-M3MiniMax
15.3%
$0.30 / $1.20
60 Mistral Medium 3.5Mistral AI
15.1%
$1.50 / $7.50
61 Gemma 4 31BGoogle · best of 2 settings
14.8%
$0.14 / $0.40
62 Kimi K2.5Moonshot AI
14.2%
$0.57 / $2.85
63 GLM-5.1Z.ai
13.6%
$1.38 / $4.40
64 Qwen3.5-397B-A17BAlibaba
13.4%
$0.55 / $3.50
64 Grok Build 0.1xAI
13.4%
$1 / $2
64 GLM-4.6 (reasoning on)Z.ai
13.4%
$0.50 / $2
67 gpt-oss-120b (high reasoning)OpenAI · best of 2 settings
12.8%
$0.15 / $0.60
68 Grok 4.3 (high reasoning)xAI · best of 2 settings
12.4%
$1.25 / $2.50
69 GLM-4.7Z.ai
12.2%
$0.54 / $1.98
70 Gemma 4 26B A4BGoogle
12.0%
$0.10 / $0.30
71 Qwen3.7-MaxAlibaba
11.8%
$1.48 / $4.42
71 GPT-5.5 Instant (2026-06-26)OpenAI
11.8%
–
73 GPT-5.2 (no reasoning)OpenAI
11.1%
$1.75 / $14
74 Devstral Small 2Mistral AI
10.7%
–
75 Devstral 2Mistral AI
10.5%
$0.40 / $2
76 Nemotron 3 Super 120B A12BNVIDIA
10.3%
$0.085 / $0.40
77 MiniMax-M2.7MiniMax
9.9%
$0.30 / $1.20
77 MiMo-V2.5-Pro (reasoning on)Xiaomi
9.9%
$0.43 / $0.87
79 Gemini 2.5 ProGoogle
9.7%
$1.25 / $10
79 Gemini 3.1 Flash-Lite PreviewGoogle
9.7%
$0.25 / $1.50
81 Mercury 2Inception
9.5%
$0.25 / $0.75
82 Qwen3.6-35B-A3BAlibaba · best of 2 settings
9.3%
$0.10 / $1
82 Claude Haiku 4.5 (reasoning on)Anthropic
9.3%
$1 / $5
84 Solar Pro 3Upstage
8.7%
$0.15 / $0.60
84 MiMo-V2.5Xiaomi
8.7%
$0.17 / $0.34
86 Mistral Medium 3.1Mistral AI
8.2%
$0.40 / $2
87 Qwen3-235B-A22B-Thinking-2507Alibaba
7.8%
$0.30 / $3
88 Granite 4.2 8BIBM
7.6%
$0.06 / $0.25
89 Mistral Small 3.1 24BMistral AI
7.4%
$0.35 / $0.56
90 Qwen3.5-9BAlibaba
7.0%
$0.10 / $0.15
90 gpt-oss-20b (high reasoning)OpenAI
7.0%
$0.03 / $0.15
92 Qwen3.5-4B (reasoning on)Alibaba · best of 2 settings
6.8%
–
93 Ministral 3 14BMistral AI
6.6%
$0.20 / $0.20
94 DeepSeek-R1DeepSeek
6.4%
$0.70 / $2.50
95 Qwen3-Next-80B-A3B-ThinkingAlibaba
6.2%
$0.15 / $1.20
95 Mistral Small 3.2 24BMistral AI
6.2%
$0.094 / $0.25
97 Nemotron 3 Nano 30B A3B (reasoning on)NVIDIA
6.0%
$0.05 / $0.20
98 Trinity Large ThinkingArcee AI
5.8%
$0.25 / $0.80
98 Mistral Large 3Mistral AI
5.8%
$0.50 / $1.50
100 Qwen3-14B (reasoning on)Alibaba
5.6%
$0.12 / $0.24
100 Magistral Medium 1.2Mistral AI
5.6%
–
102 Qwen3-30B-A3B-Thinking-2507Alibaba
5.4%
$0.20 / $2.40
102 Qwen3-32B (reasoning on)Alibaba
5.4%
$0.14 / $0.40
102 Qwen3-Coder-NextAlibaba
5.4%
$0.18 / $0.90
102 GPT-4.1 miniOpenAI
5.4%
$0.40 / $1.60
106 o3-mini (high reasoning)OpenAI
5.2%
$1.10 / $4.40
107 Qwen3.5-35B-A3B (no reasoning)Alibaba
4.9%
$0.16 / $1.30
107 Mistral Small 4Mistral AI
4.9%
$0.15 / $0.60
109 Qwen3-8B (reasoning on)Alibaba
4.7%
$0.12 / $0.46
109 DeepSeek-V3-0324DeepSeek
4.7%
$0.25 / $1
109 DeepSeek-V3DeepSeek
4.7%
$0.26 / $1.03
109 Ministral 3 3BMistral AI
4.7%
$0.10 / $0.10
113 Magistral Small 1.2Mistral AI
4.5%
–
114 Llama 4 MaverickMeta
3.7%
$0.27 / $0.85
114 Ministral 3 8BMistral AI
3.7%
$0.15 / $0.15
116 GPT-4.1 nanoOpenAI
3.5%
$0.10 / $0.40
117 Llama 4 ScoutMeta
3.3%
$0.18 / $0.59
118 GPT-4o mini (2024-07-18)OpenAI
2.9%
$0.15 / $0.60
119 Gemma 3 12BGoogle
0.8%
$0.05 / $0.15
119 Gemma 3 27BGoogle
0.8%
$0.12 / $0.20
121 Gemma 3 4BGoogle
0.4%
$0.05 / $0.10

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Results as published by Artificial Analysis; we do not re-run them.

What it measures

Share of simulated bank customer-service tasks an agent resolves with tools, as run by Artificial Analysis with its own harness and prompts.

What it does not measure

Not your policies or systems; a simulated customer.

282 results from Artificial Analysis not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Not on sale through the API providers we track: 268
  • A different snapshot or variant from the model we list: 13
  • An unusual combination of settings: 1

Data sourced from Artificial Analysis. Licence: Artificial Analysis commercial data licence.