Benchmarks / Artificial Analysis

Reported by Artificial Analysis

Artificial Analysis

Share of simulated telecom support tasks resolved with tools, as run by Artificial Analysis with its own harness and prompts.

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Results
286 configurations of 193 models
Unit
% of tasks
Licence
Artificial Analysis commercial data licence

τ²-bench telecom: Devstral Medium

Top 15 of 193 results · % of tasks, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 GLM-5.2 (max reasoning)Z.ai 99.1%
  2. 2 GLM-4.7-FlashZ.ai 98.8%
  3. 3 Claude Fable 5 (max reasoning)Anthropic 98.5%
  4. 3 GLM-5-TurboZ.ai 98.5%
  5. 3 GLM-5V-TurboZ.ai 98.5%
  6. 6 GLM-5Z.ai 98.2%
  7. 7 Qwen3.6-PlusAlibaba 97.7%
  8. 7 Grok 4.3 (high reasoning)xAI 97.7%
  9. 7 GLM-5.1Z.ai 97.7%
  10. 10 Grok 4.20 (0309, reasoning)xAI 96.5%
  11. 11 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek 96.2%
  12. 12 Qwen3.6-Max-PreviewAlibaba 95.9%
  13. 12 Kimi K2.5Moonshot AI 95.9%
  14. 12 Kimi K2.6Moonshot AI 95.9%
  15. 12 GLM-4.7Z.ai 95.9%
  16. 166 Devstral MediumMistral AI 19.9%

Full results

Artificial Analysis: τ²-bench telecom, % of tasks, higher is better
#Modelτ²-bench telecom
% of tasks, higher is better
Price
$ per million tokens, in / out
1 GLM-5.2 (max reasoning)Z.ai
99.1%
$1.40 / $4.40
2 GLM-4.7-FlashZ.ai · best of 2 settings
98.8%
$0.06 / $0.40
3 Claude Fable 5 (max reasoning)Anthropic
98.5%
$10 / $50
3 GLM-5-TurboZ.ai
98.5%
$1.20 / $4
3 GLM-5V-TurboZ.ai
98.5%
$1.20 / $4
6 GLM-5Z.ai · best of 2 settings
98.2%
$0.95 / $2.55
7 Qwen3.6-PlusAlibaba
97.7%
$0.33 / $1.95
7 Grok 4.3 (high reasoning)xAI · best of 4 settings
97.7%
$1.25 / $2.50
7 GLM-5.1Z.ai · best of 2 settings
97.7%
$1.38 / $4.40
10 Grok 4.20 (0309, reasoning)xAI
96.5%
–
11 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek · best of 3 settings
96.2%
$1.42 / $2.83
12 Qwen3.6-Max-PreviewAlibaba
95.9%
$1.03 / $6.16
12 Kimi K2.5Moonshot AI · best of 2 settings
95.9%
$0.57 / $2.85
12 Kimi K2.6Moonshot AI · best of 2 settings
95.9%
$0.95 / $4
12 GLM-4.7Z.ai · best of 2 settings
95.9%
$0.54 / $1.98
16 Qwen3.5-397B-A17BAlibaba · best of 2 settings
95.6%
$0.55 / $3.50
16 DeepSeek-V4-Flash (0423, high reasoning)DeepSeek · best of 3 settings
95.6%
$0.14 / $0.28
16 Gemini 3.1 Pro PreviewGoogle
95.6%
$2 / $12
16 Gemini 3.5 Flash (medium reasoning)Google · best of 3 settings
95.6%
$1.50 / $9
20 Qwen3.6-35B-A3BAlibaba · best of 2 settings
95.3%
$0.10 / $1
20 MiniMax-M2.5MiniMax
95.3%
$0.30 / $1.20
22 Qwen3.7-MaxAlibaba
94.7%
$1.48 / $4.42
23 Claude Opus 4.8 (max reasoning)Anthropic
94.4%
$5 / $25
23 Step 3.5 FlashStepFun
94.4%
$0.10 / $0.30
25 Qwen3.6-27BAlibaba · best of 2 settings
94.2%
$0.30 / $3.20
25 Mistral Medium 3.5Mistral AI
94.2%
$1.50 / $7.50
25 MiMo-V2.5-Pro (reasoning on)Xiaomi · best of 2 settings
94.2%
$0.43 / $0.87
28 Qwen3.5-27BAlibaba · best of 2 settings
93.9%
$0.27 / $2.16
28 GPT-5.5 (extra-high reasoning)OpenAI · best of 5 settings
93.9%
$5 / $30
30 Qwen3.5-122B-A10BAlibaba · best of 2 settings
93.6%
$0.26 / $2.08
31 Grok 4.1 Fast (reasoning)xAI
93.3%
–
32 Qwen3.7-PlusAlibaba
93.0%
$0.32 / $1.28
32 Kimi K2 ThinkingMoonshot AI
93.0%
$0.60 / $2.50
32 Grok 4.20xAI · best of 2 settings
93.0%
$1.25 / $2.50
35 Hy3 Preview (reasoning on)Tencent · best of 2 settings
92.7%
$0.18 / $0.60
36 Qwen3.5-4B (reasoning on)Alibaba · best of 2 settings
92.1%
–
36 Claude Opus 4.6 (max reasoning)Anthropic · best of 2 settings
92.1%
$5 / $25
36 GPT-5.2-Codex (extra-high reasoning)OpenAI
92.1%
$1.75 / $14
39 Muse SparkMeta
91.5%
–
40 DeepSeek-V3.2 (reasoning on)DeepSeek · best of 2 settings
90.6%
$0.30 / $0.96
40 MiMo-V2.5Xiaomi
90.6%
$0.17 / $0.34
42 Trinity Large ThinkingArcee AI
90.1%
$0.25 / $0.80
42 Kimi K2.7 CodeMoonshot AI
90.1%
$0.95 / $4
44 Claude Opus 4.5 (reasoning on)Anthropic · best of 2 settings
89.5%
$5 / $25
45 Qwen3.5-35B-A3BAlibaba · best of 2 settings
89.2%
$0.16 / $1.30
46 MiniMax-M3MiniMax
88.9%
$0.30 / $1.20
47 Claude Opus 4.7 (max reasoning)Anthropic · best of 2 settings
88.6%
$5 / $25
48 Qwen3.5-Omni-PlusAlibaba
88.3%
–
49 Gemini 3 Pro Preview (high reasoning)Google · best of 2 settings
87.1%
–
49 GPT-5.4 (extra-high reasoning)OpenAI · best of 3 settings
87.1%
$2.50 / $15
51 Qwen3.5-9BAlibaba · best of 2 settings
86.8%
$0.10 / $0.15
51 MiniMax-M2MiniMax
86.8%
$0.30 / $1.20
51 GPT-5-Codex (high reasoning)OpenAI
86.8%
–
54 GPT-5 (medium reasoning)OpenAI · best of 4 settings
86.6%
$1.25 / $10
55 GPT-5.6 Terra (max reasoning)OpenAI · best of 5 settings
86.3%
$2 / $12
55 Solar Pro 3Upstage
86.3%
$0.15 / $0.60
57 GPT-5.3-Codex (extra-high reasoning)OpenAI
86.0%
$1.75 / $14
58 MiniMax-M2.1MiniMax
85.4%
$0.30 / $1.20
59 GPT-5.6 Sol (max reasoning)OpenAI · best of 5 settings
85.1%
$4 / $20
60 MiniMax-M2.7MiniMax
84.8%
$0.30 / $1.20
60 GPT-5.2 (extra-high reasoning)OpenAI · best of 3 settings
84.8%
$1.75 / $14
62 Qwen3.5-Omni-FlashAlibaba
84.5%
–
63 Qwen3-Max-Thinking-PreviewAlibaba
83.6%
–
63 Qwen3-Max-ThinkingAlibaba
83.6%
$0.78 / $3.90
65 GPT-5.4 mini (extra-high reasoning)OpenAI · best of 3 settings
83.3%
$0.75 / $4.50
66 GPT-5.1-Codex (high reasoning)OpenAI
83.0%
$1.25 / $10
67 GPT-5.1 (high reasoning)OpenAI · best of 2 settings
81.9%
$1.25 / $10
68 Qwen3.5-2B (no reasoning)Alibaba · best of 2 settings
81.6%
–
69 o3OpenAI
80.7%
$2 / $8
70 Gemini 3 Flash Preview (reasoning on)Google · best of 2 settings
80.4%
$0.50 / $3
71 Qwen3-Coder-NextAlibaba
79.5%
$0.18 / $0.90
71 Claude Sonnet 4.6 (no reasoning)Anthropic · best of 2 settings
79.5%
$3 / $15
73 Claude Sonnet 4.5 (reasoning on)Anthropic · best of 2 settings
78.1%
$3 / $15
74 GLM-4.6 (no reasoning)Z.ai · best of 2 settings
76.9%
$0.50 / $2
75 GPT-5.4 nano (extra-high reasoning)OpenAI · best of 3 settings
76.0%
$0.20 / $1.25
76 Nova 2 Lite (medium reasoning)Amazon · best of 4 settings
75.7%
$0.30 / $2.50
76 Grok Code Fast 1xAI
75.7%
–
78 Grok 4xAI
74.9%
–
79 Qwen3-MaxAlibaba
74.3%
$0.78 / $3.90
80 Kimi K2 (0905)Moonshot AI
73.4%
$0.60 / $2.50
81 Claude Opus 4.1 (reasoning on)Anthropic
71.4%
$15 / $75
82 GPT-5 mini (medium reasoning)OpenAI · best of 3 settings
71.1%
$0.25 / $2
83 Mercury 2Inception
70.8%
$0.25 / $0.75
84 Grok 4.20 (0309, non-reasoning, no reasoning)xAI
69.6%
–
85 Nemotron 3 Super 120B A12BNVIDIA
67.8%
$0.085 / $0.40
86 gpt-oss-120b (high reasoning)OpenAI · best of 2 settings
65.8%
$0.15 / $0.60
86 Grok 4 Fast (reasoning)xAI
65.8%
–
88 Gemma 4 31B (no reasoning)Google · best of 2 settings
65.5%
$0.14 / $0.40
89 Qwen3.5-0.8B (no reasoning)Alibaba · best of 2 settings
65.2%
–
90 Claude Sonnet 4 (reasoning on)Anthropic · best of 2 settings
64.6%
$3 / $15
91 Grok 4.1 Fast (non-reasoning, no reasoning)xAI
63.7%
–
91 Grok 4 Fast (non-reasoning, no reasoning)xAI
63.7%
–
93 GPT-5.1-Codex-Mini (high reasoning)OpenAI
62.9%
$0.25 / $2
94 o1OpenAI
62.6%
$15 / $60
95 Kimi K2 (0711)Moonshot AI
61.1%
$0.57 / $2.30
96 gpt-oss-20b (high reasoning)OpenAI · best of 2 settings
60.2%
$0.03 / $0.15
97 o4-mini (high reasoning)OpenAI
55.6%
$1.10 / $4.40
98 Claude Haiku 4.5 (reasoning on)Anthropic · best of 2 settings
54.7%
$1 / $5
99 Qwen3-VL-235B-A22B-ThinkingAlibaba
54.1%
$0.40 / $4
99 Gemini 2.5 ProGoogle
54.1%
$1.25 / $10
101 Qwen3-235B-A22B-Thinking-2507Alibaba
53.2%
$0.30 / $3
102 GPT-4.1 miniOpenAI
52.9%
$0.40 / $1.60
103 Magistral Medium 1.2Mistral AI
52.0%
–
104 GPT-5.5 Instant (2026-05-26)OpenAI
49.4%
–
105 DeepSeek-V3-0324DeepSeek
47.1%
$0.25 / $1
105 GPT-4.1OpenAI
47.1%
$2 / $8
107 GLM-4.5-AirZ.ai
46.5%
$0.14 / $0.86
108 Qwen3-VL-32B-ThinkingAlibaba
45.6%
–
108 Gemini 2.5 Flash Preview (09-2025, reasoning on)Google · best of 2 settings
45.6%
–
110 Qwen3-Coder-480B-A35BAlibaba
43.6%
$0.35 / $1.50
110 Gemma 4 26B A4BGoogle · best of 2 settings
43.6%
$0.10 / $0.30
112 GLM-4.5Z.ai
43.0%
$0.60 / $2.20
113 Qwen3-Next-80B-A3B-ThinkingAlibaba
41.5%
$0.15 / $1.20
114 Mistral Small 4Mistral AI · best of 2 settings
41.2%
$0.15 / $0.60
115 Nemotron 3 Nano 30B A3B (reasoning on)NVIDIA · best of 2 settings
40.9%
$0.05 / $0.20
116 Mistral Medium 3.1Mistral AI
40.6%
$0.40 / $2
117 DeepSeek-V3.1 (reasoning on)DeepSeek · best of 2 settings
37.4%
$0.55 / $1.65
118 DeepSeek-V3.1-Terminus (no reasoning)DeepSeek · best of 2 settings
37.1%
$0.27 / $1
119 DeepSeek-R1-0528DeepSeek
36.6%
$0.50 / $2.18
119 GPT-5 nano (high reasoning)OpenAI · best of 3 settings
36.6%
$0.05 / $0.40
121 Gemma 4 12BGoogle · best of 2 settings
36.3%
–
122 Qwen3-VL-235B-A22B-InstructAlibaba
35.1%
$0.30 / $1.50
123 Qwen2.5-72B-InstructAlibaba
34.5%
$0.36 / $0.40
123 Qwen3-14B (reasoning on)Alibaba · best of 2 settings
34.5%
$0.12 / $0.24
123 Qwen3-Coder-30B-A3B-InstructAlibaba
34.5%
$0.07 / $0.28
126 DeepSeek-V3.2-Exp (no reasoning)DeepSeek · best of 2 settings
33.9%
$0.27 / $0.41
127 Qwen3-235B-A22B-Instruct-2507Alibaba
33.3%
$0.15 / $0.75
128 Mistral Large 2 (2407)Mistral AI
33.0%
$2 / $6
129 Qwen3-Max-PreviewAlibaba
32.7%
–
130 Gemini 2.5 Flash (reasoning on)Google · best of 2 settings
31.6%
$0.30 / $2.50
130 GLM-4.6V (reasoning on)Z.ai · best of 2 settings
31.6%
$0.30 / $0.90
132 Gemini 3.1 Flash-Lite PreviewGoogle
31.3%
$0.25 / $1.50
132 o3-mini (high reasoning)OpenAI · best of 2 settings
31.3%
$1.10 / $4.40
134 Gemini 2.5 Flash-Lite Preview (09-2025, reasoning on)Google · best of 2 settings
30.7%
–
135 Qwen3-32B (reasoning on)Alibaba
29.8%
$0.14 / $0.40
136 Mistral Small 3.2 24BMistral AI
29.5%
$0.094 / $0.25
137 Qwen3-VL-32B-InstructAlibaba
29.2%
$0.10 / $0.42
137 Qwen3-VL-8B-InstructAlibaba
29.2%
$0.12 / $0.46
139 GPT-4o (2024-08-06)OpenAI
28.9%
$2.50 / $10
140 Devstral Small 1.1Mistral AI
28.4%
–
141 Qwen3-30B-A3B-Thinking-2507Alibaba
28.1%
$0.20 / $2.40
142 Qwen3-8B (reasoning on)Alibaba · best of 2 settings
27.8%
$0.12 / $0.46
142 Magistral Small 1.2Mistral AI
27.8%
–
144 Qwen3-235B-A22B (no reasoning)Alibaba · best of 2 settings
27.2%
$0.46 / $1.82
144 Ministral 3 14BMistral AI
27.2%
$0.20 / $0.20
146 Qwen3-4B-Instruct-2507Alibaba
26.6%
–
146 Llama 3.3 70B InstructMeta
26.6%
$0.59 / $0.79
146 Ministral 3 8BMistral AI
26.6%
$0.15 / $0.15
149 Qwen3-30B-A3B (reasoning on)Alibaba · best of 2 settings
26.0%
$0.12 / $0.50
149 Gemma 4 E4B (no reasoning)Google · best of 2 settings
26.0%
–
151 Qwen3-4B-Thinking-2507Alibaba
25.4%
–
152 Mistral Small 3.1 24BMistral AI
25.1%
$0.35 / $0.56
153 Devstral 2Mistral AI
24.9%
$0.40 / $2
153 Ministral 3 3BMistral AI
24.9%
$0.10 / $0.10
155 Mistral Large 3Mistral AI
24.6%
$0.50 / $1.50
156 Mistral Medium 3Mistral AI
24.3%
$0.40 / $2
157 Qwen3-VL-4B-InstructAlibaba
23.4%
–
157 Devstral Small 2Mistral AI
23.4%
–
159 DeepSeek-V3DeepSeek
22.8%
$0.26 / $1.03
160 Qwen3-VL-8B-ThinkingAlibaba
22.5%
$0.18 / $2.10
160 GLM-4.5V (reasoning on)Z.ai · best of 2 settings
22.5%
$0.60 / $1.80
162 Gemma 4 E2B (no reasoning)Google · best of 2 settings
22.2%
–
163 Qwen3-Next-80B-A3B-InstructAlibaba
21.6%
$0.10 / $1.10
164 Qwen3-Omni-30B-A3B-ThinkingAlibaba
21.3%
–
165 Llama 3.2 3B InstructMeta
21.1%
$0.05 / $0.33
166 Qwen3-VL-30B-A3B-ThinkingAlibaba
19.9%
$0.29 / $1
166 Devstral MediumMistral AI
19.9%
–
168 Mistral Small 3Mistral AI
19.6%
$0.05 / $0.08
169 Qwen3-VL-30B-A3B-InstructAlibaba
19.0%
$0.15 / $0.60
169 Gemini 2.5 Flash-Lite (no reasoning)Google · best of 2 settings
19.0%
$0.10 / $0.40
171 Llama 4 MaverickMeta
17.8%
$0.27 / $0.85
172 Nova Lite 1.0Amazon
17.5%
$0.06 / $0.24
173 GPT-4.1 nanoOpenAI
17.3%
$0.10 / $0.40
174 Qwen3-Omni-30B-A3B-InstructAlibaba
16.4%
–
174 Llama 3.1 8B InstructMeta
16.4%
$0.05 / $0.08
176 Qwen3-VL-4B-ThinkingAlibaba
15.5%
–
176 Llama 4 ScoutMeta
15.5%
$0.18 / $0.59
178 Command ACohere
15.2%
$2.50 / $10
178 Llama 3.1 70B InstructMeta
15.2%
$0.40 / $0.40
180 Nova Micro 1.0Amazon
14.0%
$0.035 / $0.14
180 Nova Pro 1.0Amazon
14.0%
$0.80 / $3.20
182 DeepSeek-R1DeepSeek
11.4%
$0.70 / $2.50
183 Gemma 3 12BGoogle
10.8%
$0.05 / $0.15
184 Gemma 3 27BGoogle
10.5%
$0.12 / $0.20
185 Qwen3-30B-A3B-Instruct-2507Alibaba
10.2%
$0.09 / $0.30
186 Gemma 3 270MGoogle
9.1%
–
187 Gemma 3 4BGoogle
5.0%
$0.05 / $0.10
188 DeepSeek-V3.2-SpecialeDeepSeek
0.0%
–
188 Llama 3.2 1B InstructMeta
0.0%
$0.027 / $0.20
188 Phi-4Microsoft
0.0%
$0.07 / $0.14
188 Kimi Linear 48B A3B InstructMoonshot AI
0.0%
–
188 GPT-5 ChatOpenAI
0.0%
–
188 Reka Flash 3Reka AI
0.0%
$0.10 / $0.20

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Results as published by Artificial Analysis; we do not re-run them.

What it measures

Share of simulated telecom support tasks resolved with tools, as run by Artificial Analysis with its own harness and prompts.

What it does not measure

Not your policies or systems; no longer run on new models.

282 results from Artificial Analysis not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Not on sale through the API providers we track: 268
  • A different snapshot or variant from the model we list: 13
  • An unusual combination of settings: 1

Data sourced from Artificial Analysis. Licence: Artificial Analysis commercial data licence.