Benchmarks / Artificial Analysis

Reported by Artificial Analysis

Artificial Analysis

Share of simulated bank customer-service tasks an agent resolves with tools, as run by Artificial Analysis with its own harness and prompts.

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Results
175 configurations of 121 models
Unit
% of tasks
Licence
Artificial Analysis commercial data licence

τ-bench banking: GLM-5.3-Flash

Top 15 of 175 results · % of tasks, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 Qwen3.8-Max (0803)Alibaba 51.3%
  2. 2 Grok 4.6 (high reasoning)xAI 50.7%
  3. 3 Muse Spark 1.3 (max reasoning)Meta 50.5%
  4. 4 GLM-5.3 (max reasoning)Z.ai 50.3%
  5. 5 Qwen3.8-2.4T-A95BAlibaba 49.1%
  6. 6 Qwen3.8-27B (extra-high reasoning)Alibaba 48.0%
  7. 7 Qwen3.8-Max (0902)Alibaba 47.8%
  8. 8 Qwen3.8-27B (medium reasoning)Alibaba 47.4%
  9. 9 Claude Fable 5.1 (max reasoning)Anthropic 47.2%
  10. 9 Muse Spark 1.3 (extra-high reasoning)Meta 47.2%
  11. 9 GLM-5.3-FlashZ.ai 47.2%
  12. 12 Kimi K3 (max reasoning)Moonshot AI 46.0%
  13. 13 Claude Fable 5.1 (extra-high reasoning)Anthropic 45.8%
  14. 13 Gemini 3.8 Flash (medium reasoning)Google 45.8%
  15. 15 Qwen3.8-Flash-NextAlibaba 45.4%

Full results

Artificial Analysis: τ-bench banking, % of tasks, higher is better
#Modelτ-bench banking
% of tasks, higher is better
Price
$ per million tokens, in / out
1 Qwen3.8-Max (0803)Alibaba
51.3%
–
2 Grok 4.6 (high reasoning)xAI
50.7%
$2 / $6
3 Muse Spark 1.3 (max reasoning)Meta
50.5%
$1.25 / $4.25
4 GLM-5.3 (max reasoning)Z.ai
50.3%
$1.40 / $4.40
5 Qwen3.8-2.4T-A95BAlibaba
49.1%
$2 / $6
6 Qwen3.8-27B (extra-high reasoning)Alibaba
48.0%
$0.50 / $3
7 Qwen3.8-Max (0902)Alibaba
47.8%
$2 / $6
8 Qwen3.8-27B (medium reasoning)Alibaba
47.4%
$0.50 / $3
9 Claude Fable 5.1 (max reasoning)Anthropic
47.2%
$10 / $50
9 Muse Spark 1.3 (extra-high reasoning)Meta
47.2%
$1.25 / $4.25
9 GLM-5.3-FlashZ.ai
47.2%
$0.15 / $0.50
12 Kimi K3 (max reasoning)Moonshot AI
46.0%
$3 / $15
13 Claude Fable 5.1 (extra-high reasoning)Anthropic
45.8%
$10 / $50
13 Gemini 3.8 Flash (medium reasoning)Google
45.8%
$1.50 / $7.50
15 Qwen3.8-Flash-NextAlibaba
45.4%
–
16 Gemini 3.8 Flash (high reasoning)Google
44.9%
$1.50 / $7.50
17 Claude Opus 5 (high reasoning)Anthropic
44.7%
$5 / $25
18 GPT-5.6 Sol (max reasoning)OpenAI
44.3%
$4 / $20
18 Grok 4.6 (medium reasoning)xAI
44.3%
$2 / $6
20 Claude Opus 5 (extra-high reasoning)Anthropic
43.3%
$5 / $25
20 Grok 4.6 (extra-high reasoning)xAI
43.3%
$2 / $6
22 Claude Fable 5.1 (high reasoning)Anthropic
43.1%
$10 / $50
22 GPT-6 Astra (extra-high reasoning)OpenAI
43.1%
$10 / $50
24 Claude Opus 5 (max reasoning)Anthropic
42.1%
$5 / $25
24 Grok 4.5 (high reasoning)xAI
42.1%
$2 / $6
26 Kimi K3 (low reasoning)Moonshot AI
41.6%
$3 / $15
27 GPT-6 Astra (max reasoning)OpenAI
41.4%
$10 / $50
28 Claude Fable 5.1 (medium reasoning)Anthropic
41.0%
$10 / $50
28 DeepSeek-V4-Flash-Vision-Exp (max reasoning)DeepSeek
41.0%
$0.44 / $1.32
30 GPT-5.6 Terra (max reasoning)OpenAI
40.2%
$2 / $12
31 GPT-6 Astra (high reasoning)OpenAI
40.0%
$10 / $50
32 DeepSeek-V4-Pro (0813, max reasoning)DeepSeek
39.6%
$1.32 / $3.96
32 GPT-5.4 (extra-high reasoning)OpenAI
39.6%
$2.50 / $15
34 DeepSeek-V4-Flash (0731, max reasoning)DeepSeek
39.4%
$0.14 / $0.28
35 Claude Fable 5.1 (low reasoning)Anthropic
39.0%
$10 / $50
35 GPT-5.5 (extra-high reasoning)OpenAI
39.0%
$5 / $30
37 Claude Opus 5 (medium reasoning)Anthropic
38.6%
$5 / $25
38 Claude Fable 5 (max reasoning)Anthropic
38.1%
$10 / $50
38 GPT-5.6 Sol (extra-high reasoning)OpenAI
38.1%
$4 / $20
38 Grok 4.6 (low reasoning)xAI
38.1%
$2 / $6
41 Claude Sonnet 5 (max reasoning)Anthropic
37.3%
$2 / $10
42 GPT-5.5 (high reasoning)OpenAI
36.7%
$5 / $30
42 GPT-5.6 Sol (high reasoning)OpenAI
36.7%
$4 / $20
44 GPT-5.6 Sol (medium reasoning)OpenAI
36.5%
$4 / $20
45 Gemini 3.7 Flash (medium reasoning)Google
35.5%
$1.50 / $7.50
45 GPT-6 Astra (medium reasoning)OpenAI
35.5%
$10 / $50
47 Muse Spark 1.2 (extra-high reasoning)Meta
34.8%
$1.25 / $4.25
48 Claude Opus 4.7 (max reasoning)Anthropic
34.6%
$5 / $25
48 GLM-5.2 (max reasoning)Z.ai
34.6%
$1.40 / $4.40
50 Claude Sonnet 4.6 (max reasoning)Anthropic
34.4%
$3 / $15
51 Claude Opus 4.8 (max reasoning)Anthropic
34.2%
$5 / $25
52 Gemini 3.8 Flash (low reasoning)Google
33.2%
$1.50 / $7.50
53 Gemini 3.7 Flash (high reasoning)Google
32.8%
$1.50 / $7.50
54 Qwen3.8-27B (low reasoning)Alibaba
32.2%
$0.50 / $3
54 Gemini 3.5 Flash (high reasoning)Google
32.2%
$1.50 / $9
56 GPT-6 Astra (low reasoning)OpenAI
32.0%
$10 / $50
57 Muse Spark 1.1 (extra-high reasoning)Meta
31.8%
$1.25 / $4.25
58 GPT-5.6 Luna (max reasoning)OpenAI
31.1%
$0.20 / $1.20
59 DeepSeek-V4-Flash (0423, max reasoning)DeepSeek
30.9%
$0.14 / $0.28
60 Claude Opus 5 (low reasoning)Anthropic
30.3%
$5 / $25
61 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek
30.1%
$1.42 / $2.83
62 Gemini 3.6 Flash (high reasoning)Google
29.9%
$1.50 / $7.50
62 GPT-5.5 (medium reasoning)OpenAI
29.9%
$5 / $30
64 GPT-5.6 Terra (extra-high reasoning)OpenAI
29.7%
$2 / $12
65 Gemini 3.7 Flash (low reasoning)Google
29.5%
$1.50 / $7.50
66 GPT-5.6 Sol (low reasoning)OpenAI
29.1%
$4 / $20
66 Inkling (extra-high reasoning)Thinking Machines
29.1%
$0.95 / $4.05
68 GPT-5.6 Luna (extra-high reasoning)OpenAI
28.7%
$0.20 / $1.20
68 GPT-5.6 Terra (high reasoning)OpenAI
28.7%
$2 / $12
70 GPT-5.4 nano (extra-high reasoning)OpenAI
27.4%
$0.20 / $1.25
71 DeepSeek-V4-Flash (0423, high reasoning)DeepSeek
26.2%
$0.14 / $0.28
71 DeepSeek-V4-Pro (0423, high reasoning)DeepSeek
26.2%
$1.42 / $2.83
73 GPT-5.4 mini (extra-high reasoning)OpenAI
25.6%
$0.75 / $4.50
73 GPT-5.6 Terra (medium reasoning)OpenAI
25.6%
$2 / $12
75 GPT-5.6 Luna (high reasoning)OpenAI
25.2%
$0.20 / $1.20
76 GPT-5.5 (low reasoning)OpenAI
24.9%
$5 / $30
77 Claude Sonnet 4.5 (reasoning on)Anthropic
24.5%
$3 / $15
78 Muse Glimmer (high reasoning)Meta
23.5%
–
79 Kimi K2.6Moonshot AI
23.3%
$0.95 / $4
79 Solar Pro 4Upstage
23.3%
$0.09 / $0.36
81 Hy3Tencent
22.9%
$0.14 / $0.58
82 GPT-5 (high reasoning)OpenAI
22.1%
$1.25 / $10
83 Gemini 3.1 Pro PreviewGoogle
21.4%
$2 / $12
84 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek
21.0%
$0.27 / $1
85 Qwen3.6-PlusAlibaba
20.8%
$0.33 / $1.95
85 Gemini 3 Flash Preview (reasoning on)Google
20.8%
$0.50 / $3
87 Kimi K2.7 CodeMoonshot AI
20.2%
$0.95 / $4
88 Qwen3.8-27B (no reasoning)Alibaba
20.0%
$0.50 / $3
89 GPT-5.6 Sol (no reasoning)OpenAI
19.6%
$4 / $20
90 GPT-5.6 Terra (low reasoning)OpenAI
18.8%
$2 / $12
90 Inkling SmallThinking Machines
18.8%
$0.45 / $1.20
92 GPT-5.6 Luna (medium reasoning)OpenAI
17.7%
$0.20 / $1.20
93 Qwen3.7-PlusAlibaba
17.5%
$0.32 / $1.28
93 Gemini 3.5 Flash-LiteGoogle
17.5%
$0.30 / $2.50
95 Qwen3.6-27BAlibaba
16.7%
$0.30 / $3.20
95 Claude Sonnet 4 (reasoning on)Anthropic
16.7%
$3 / $15
95 GLM-5.2 (no reasoning)Z.ai
16.7%
$1.40 / $4.40
98 GPT-5.1 (high reasoning)OpenAI
15.9%
$1.25 / $10
99 Claude Sonnet 5 (no reasoning)Anthropic
15.7%
$2 / $10
99 GPT-5.6 Terra (no reasoning)OpenAI
15.7%
$2 / $12
101 GPT-5 mini (high reasoning)OpenAI
15.5%
$0.25 / $2
102 Qwen3.5-122B-A10BAlibaba
15.3%
$0.26 / $2.08
102 MiniMax-M3MiniMax
15.3%
$0.30 / $1.20
104 Mistral Medium 3.5Mistral AI
15.1%
$1.50 / $7.50
105 Gemma 4 31BGoogle
14.8%
$0.14 / $0.40
105 GPT-5.5 (no reasoning)OpenAI
14.8%
$5 / $30
107 Kimi K2.5Moonshot AI
14.2%
$0.57 / $2.85
108 GLM-5.1Z.ai
13.6%
$1.38 / $4.40
109 Qwen3.5-397B-A17BAlibaba
13.4%
$0.55 / $3.50
109 Grok Build 0.1xAI
13.4%
$1 / $2
109 GLM-4.6 (reasoning on)Z.ai
13.4%
$0.50 / $2
112 GPT-5.6 Luna (low reasoning)OpenAI
12.8%
$0.20 / $1.20
112 gpt-oss-120b (high reasoning)OpenAI
12.8%
$0.15 / $0.60
114 Grok 4.3 (high reasoning)xAI
12.4%
$1.25 / $2.50
115 GLM-4.7Z.ai
12.2%
$0.54 / $1.98
116 Gemma 4 26B A4BGoogle
12.0%
$0.10 / $0.30
117 Qwen3.7-MaxAlibaba
11.8%
$1.48 / $4.42
117 GPT-5.5 Instant (2026-06-26)OpenAI
11.8%
–
119 GPT-5.2 (no reasoning)OpenAI
11.1%
$1.75 / $14
120 Devstral Small 2Mistral AI
10.7%
–
121 Devstral 2Mistral AI
10.5%
$0.40 / $2
122 Qwen3.5-122B-A10B (no reasoning)Alibaba
10.3%
$0.26 / $2.08
122 Nemotron 3 Super 120B A12BNVIDIA
10.3%
$0.085 / $0.40
124 MiniMax-M2.7MiniMax
9.9%
$0.30 / $1.20
124 MiMo-V2.5-Pro (reasoning on)Xiaomi
9.9%
$0.43 / $0.87
126 Gemini 2.5 ProGoogle
9.7%
$1.25 / $10
126 Gemini 3.1 Flash-Lite PreviewGoogle
9.7%
$0.25 / $1.50
126 GPT-5.6 Luna (no reasoning)OpenAI
9.7%
$0.20 / $1.20
129 Mercury 2Inception
9.5%
$0.25 / $0.75
130 Qwen3.6-27B (no reasoning)Alibaba
9.3%
$0.30 / $3.20
130 Qwen3.6-35B-A3BAlibaba
9.3%
$0.10 / $1
130 Claude Haiku 4.5 (reasoning on)Anthropic
9.3%
$1 / $5
133 Gemma 4 31B (no reasoning)Google
8.9%
$0.14 / $0.40
134 Solar Pro 3Upstage
8.7%
$0.15 / $0.60
134 MiMo-V2.5Xiaomi
8.7%
$0.17 / $0.34
136 Mistral Medium 3.1Mistral AI
8.2%
$0.40 / $2
137 Grok 4.3 (no reasoning)xAI
8.0%
$1.25 / $2.50
138 Qwen3-235B-A22B-Thinking-2507Alibaba
7.8%
$0.30 / $3
139 Granite 4.2 8BIBM
7.6%
$0.06 / $0.25
140 Mistral Small 3.1 24BMistral AI
7.4%
$0.35 / $0.56
141 Qwen3.5-9BAlibaba
7.0%
$0.10 / $0.15
141 gpt-oss-20b (high reasoning)OpenAI
7.0%
$0.03 / $0.15
143 Qwen3.5-4B (reasoning on)Alibaba
6.8%
–
144 Ministral 3 14BMistral AI
6.6%
$0.20 / $0.20
145 DeepSeek-R1DeepSeek
6.4%
$0.70 / $2.50
146 Qwen3-Next-80B-A3B-ThinkingAlibaba
6.2%
$0.15 / $1.20
146 Mistral Small 3.2 24BMistral AI
6.2%
$0.094 / $0.25
148 Nemotron 3 Nano 30B A3B (reasoning on)NVIDIA
6.0%
$0.05 / $0.20
149 Trinity Large ThinkingArcee AI
5.8%
$0.25 / $0.80
149 Mistral Large 3Mistral AI
5.8%
$0.50 / $1.50
151 Qwen3-14B (reasoning on)Alibaba
5.6%
$0.12 / $0.24
151 Magistral Medium 1.2Mistral AI
5.6%
–
153 Qwen3-30B-A3B-Thinking-2507Alibaba
5.4%
$0.20 / $2.40
153 Qwen3-32B (reasoning on)Alibaba
5.4%
$0.14 / $0.40
153 Qwen3.6-35B-A3B (no reasoning)Alibaba
5.4%
$0.10 / $1
153 Qwen3-Coder-NextAlibaba
5.4%
$0.18 / $0.90
153 GPT-4.1 miniOpenAI
5.4%
$0.40 / $1.60
158 o3-mini (high reasoning)OpenAI
5.2%
$1.10 / $4.40
159 Qwen3.5-35B-A3B (no reasoning)Alibaba
4.9%
$0.16 / $1.30
159 Mistral Small 4Mistral AI
4.9%
$0.15 / $0.60
161 Qwen3-8B (reasoning on)Alibaba
4.7%
$0.12 / $0.46
161 DeepSeek-V3-0324DeepSeek
4.7%
$0.25 / $1
161 DeepSeek-V3DeepSeek
4.7%
$0.26 / $1.03
161 Ministral 3 3BMistral AI
4.7%
$0.10 / $0.10
165 Magistral Small 1.2Mistral AI
4.5%
–
166 Qwen3.5-4B (no reasoning)Alibaba
4.3%
–
167 Llama 4 MaverickMeta
3.7%
$0.27 / $0.85
167 Ministral 3 8BMistral AI
3.7%
$0.15 / $0.15
169 GPT-4.1 nanoOpenAI
3.5%
$0.10 / $0.40
170 Llama 4 ScoutMeta
3.3%
$0.18 / $0.59
171 GPT-4o mini (2024-07-18)OpenAI
2.9%
$0.15 / $0.60
171 gpt-oss-120b (low reasoning)OpenAI
2.9%
$0.15 / $0.60
173 Gemma 3 12BGoogle
0.8%
$0.05 / $0.15
173 Gemma 3 27BGoogle
0.8%
$0.12 / $0.20
175 Gemma 3 4BGoogle
0.4%
$0.05 / $0.10

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Results as published by Artificial Analysis; we do not re-run them.

What it measures

Share of simulated bank customer-service tasks an agent resolves with tools, as run by Artificial Analysis with its own harness and prompts.

What it does not measure

Not your policies or systems; a simulated customer.

282 results from Artificial Analysis not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Not on sale through the API providers we track: 268
  • A different snapshot or variant from the model we list: 13
  • An unusual combination of settings: 1

Data sourced from Artificial Analysis. Licence: Artificial Analysis commercial data licence.