Benchmarks / Artificial Analysis

Reported by Artificial Analysis

Artificial Analysis

Share of simulated telecom support tasks resolved with tools, as run by Artificial Analysis with its own harness and prompts.

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Results
286 configurations of 193 models
Unit
% of tasks
Licence
Artificial Analysis commercial data licence

τ²-bench telecom: Qwen3-Omni-30B-A3B-Instruct

Top 15 of 286 results · % of tasks, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 GLM-5.2 (max reasoning)Z.ai 99.1%
  2. 2 GLM-4.7-FlashZ.ai 98.8%
  3. 3 Claude Fable 5 (max reasoning)Anthropic 98.5%
  4. 3 GLM-5-TurboZ.ai 98.5%
  5. 3 GLM-5V-TurboZ.ai 98.5%
  6. 6 GLM-5Z.ai 98.2%
  7. 7 Qwen3.6-PlusAlibaba 97.7%
  8. 7 Grok 4.3 (high reasoning)xAI 97.7%
  9. 7 GLM-5.1Z.ai 97.7%
  10. 10 GLM-5 (no reasoning)Z.ai 97.4%
  11. 11 GLM-5.1 (no reasoning)Z.ai 97.1%
  12. 12 Grok 4.20 (0309, reasoning)xAI 96.5%
  13. 13 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek 96.2%
  14. 14 Qwen3.6-Max-PreviewAlibaba 95.9%
  15. 14 Kimi K2.5Moonshot AI 95.9%
  16. 266 Qwen3-Omni-30B-A3B-InstructAlibaba 16.4%

Full results

Artificial Analysis: τ²-bench telecom, % of tasks, higher is better
#Modelτ²-bench telecom
% of tasks, higher is better
Price
$ per million tokens, in / out
1 GLM-5.2 (max reasoning)Z.ai
99.1%
$1.40 / $4.40
2 GLM-4.7-FlashZ.ai
98.8%
$0.06 / $0.40
3 Claude Fable 5 (max reasoning)Anthropic
98.5%
$10 / $50
3 GLM-5-TurboZ.ai
98.5%
$1.20 / $4
3 GLM-5V-TurboZ.ai
98.5%
$1.20 / $4
6 GLM-5Z.ai
98.2%
$0.95 / $2.55
7 Qwen3.6-PlusAlibaba
97.7%
$0.33 / $1.95
7 Grok 4.3 (high reasoning)xAI
97.7%
$1.25 / $2.50
7 GLM-5.1Z.ai
97.7%
$1.38 / $4.40
10 GLM-5 (no reasoning)Z.ai
97.4%
$0.95 / $2.55
11 GLM-5.1 (no reasoning)Z.ai
97.1%
$1.38 / $4.40
12 Grok 4.20 (0309, reasoning)xAI
96.5%
–
13 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek
96.2%
$1.42 / $2.83
14 Qwen3.6-Max-PreviewAlibaba
95.9%
$1.03 / $6.16
14 Kimi K2.5Moonshot AI
95.9%
$0.57 / $2.85
14 Kimi K2.6Moonshot AI
95.9%
$0.95 / $4
14 GLM-4.7Z.ai
95.9%
$0.54 / $1.98
18 Qwen3.5-397B-A17BAlibaba
95.6%
$0.55 / $3.50
18 DeepSeek-V4-Flash (0423, high reasoning)DeepSeek
95.6%
$0.14 / $0.28
18 Gemini 3.1 Pro PreviewGoogle
95.6%
$2 / $12
18 Gemini 3.5 Flash (medium reasoning)Google
95.6%
$1.50 / $9
22 Qwen3.6-35B-A3BAlibaba
95.3%
$0.10 / $1
22 Gemini 3.5 Flash (high reasoning)Google
95.3%
$1.50 / $9
22 MiniMax-M2.5MiniMax
95.3%
$0.30 / $1.20
25 DeepSeek-V4-Flash (0423, max reasoning)DeepSeek
95.0%
$0.14 / $0.28
26 Qwen3.7-MaxAlibaba
94.7%
$1.48 / $4.42
27 Claude Opus 4.8 (max reasoning)Anthropic
94.4%
$5 / $25
27 DeepSeek-V4-Flash (0423, no reasoning)DeepSeek
94.4%
$0.14 / $0.28
27 Step 3.5 FlashStepFun
94.4%
$0.10 / $0.30
30 Qwen3.6-27BAlibaba
94.2%
$0.30 / $3.20
30 DeepSeek-V4-Pro (0423, high reasoning)DeepSeek
94.2%
$1.42 / $2.83
30 Mistral Medium 3.5Mistral AI
94.2%
$1.50 / $7.50
30 MiMo-V2.5-Pro (reasoning on)Xiaomi
94.2%
$0.43 / $0.87
30 GLM-4.7 (no reasoning)Z.ai
94.2%
$0.54 / $1.98
35 Qwen3.5-27BAlibaba
93.9%
$0.27 / $2.16
35 Kimi K2.6 (no reasoning)Moonshot AI
93.9%
$0.95 / $4
35 GPT-5.5 (extra-high reasoning)OpenAI
93.9%
$5 / $30
38 Qwen3.5-122B-A10BAlibaba
93.6%
$0.26 / $2.08
38 Qwen3.6-27B (no reasoning)Alibaba
93.6%
$0.30 / $3.20
40 Grok 4.1 Fast (reasoning)xAI
93.3%
–
41 Qwen3.7-PlusAlibaba
93.0%
$0.32 / $1.28
41 Kimi K2 ThinkingMoonshot AI
93.0%
$0.60 / $2.50
41 GPT-5.5 (high reasoning)OpenAI
93.0%
$5 / $30
41 Grok 4.20xAI
93.0%
$1.25 / $2.50
45 Hy3 Preview (reasoning on)Tencent
92.7%
$0.18 / $0.60
46 Qwen3.5-4B (reasoning on)Alibaba
92.1%
–
46 Claude Opus 4.6 (max reasoning)Anthropic
92.1%
$5 / $25
46 GPT-5.2-Codex (extra-high reasoning)OpenAI
92.1%
$1.75 / $14
49 GPT-5.5 (medium reasoning)OpenAI
91.8%
$5 / $30
49 GLM-4.7-Flash (no reasoning)Z.ai
91.8%
$0.06 / $0.40
51 Muse SparkMeta
91.5%
–
52 DeepSeek-V4-Pro (0423, no reasoning)DeepSeek
91.2%
$1.42 / $2.83
52 Grok 4.3 (medium reasoning)xAI
91.2%
$1.25 / $2.50
54 DeepSeek-V3.2 (reasoning on)DeepSeek
90.6%
$0.30 / $0.96
54 MiMo-V2.5Xiaomi
90.6%
$0.17 / $0.34
56 Trinity Large ThinkingArcee AI
90.1%
$0.25 / $0.80
56 Kimi K2.7 CodeMoonshot AI
90.1%
$0.95 / $4
58 Claude Opus 4.5 (reasoning on)Anthropic
89.5%
$5 / $25
59 Qwen3.5-35B-A3BAlibaba
89.2%
$0.16 / $1.30
60 MiniMax-M3MiniMax
88.9%
$0.30 / $1.20
60 Grok 4.3 (low reasoning)xAI
88.9%
$1.25 / $2.50
62 Claude Opus 4.7 (max reasoning)Anthropic
88.6%
$5 / $25
63 Qwen3.5-Omni-PlusAlibaba
88.3%
–
64 Qwen3.5-4B (no reasoning)Alibaba
87.7%
–
65 Qwen3.5-27B (no reasoning)Alibaba
87.1%
$0.27 / $2.16
65 Gemini 3 Pro Preview (high reasoning)Google
87.1%
–
65 GPT-5.4 (extra-high reasoning)OpenAI
87.1%
$2.50 / $15
68 Qwen3.5-9BAlibaba
86.8%
$0.10 / $0.15
68 MiniMax-M2MiniMax
86.8%
$0.30 / $1.20
68 GPT-5-Codex (high reasoning)OpenAI
86.8%
–
71 GPT-5 (medium reasoning)OpenAI
86.6%
$1.25 / $10
72 Qwen3.5-35B-A3B (no reasoning)Alibaba
86.3%
$0.16 / $1.30
72 Claude Opus 4.5 (no reasoning)Anthropic
86.3%
$5 / $25
72 GPT-5.6 Terra (max reasoning)OpenAI
86.3%
$2 / $12
72 Solar Pro 3Upstage
86.3%
$0.15 / $0.60
76 GPT-5.3-Codex (extra-high reasoning)OpenAI
86.0%
$1.75 / $14
77 MiniMax-M2.1MiniMax
85.4%
$0.30 / $1.20
78 Qwen3.5-9B (no reasoning)Alibaba
85.1%
$0.10 / $0.15
78 Qwen3.6-35B-A3B (no reasoning)Alibaba
85.1%
$0.10 / $1
78 GPT-5.6 Sol (max reasoning)OpenAI
85.1%
$4 / $20
81 Claude Opus 4.6 (no reasoning)Anthropic
84.8%
$5 / $25
81 MiniMax-M2.7MiniMax
84.8%
$0.30 / $1.20
81 GPT-5.2 (extra-high reasoning)OpenAI
84.8%
$1.75 / $14
81 GPT-5.6 Sol (extra-high reasoning)OpenAI
84.8%
$4 / $20
81 GPT-5 (high reasoning)OpenAI
84.8%
$1.25 / $10
86 Qwen3.5-122B-A10B (no reasoning)Alibaba
84.5%
$0.26 / $2.08
86 Qwen3.5-Omni-FlashAlibaba
84.5%
–
88 GPT-5 (low reasoning)OpenAI
84.2%
$1.25 / $10
89 Qwen3.5-397B-A17B (no reasoning)Alibaba
83.9%
$0.55 / $3.50
89 GPT-5.5 (low reasoning)OpenAI
83.9%
$5 / $30
91 Qwen3-Max-Thinking-PreviewAlibaba
83.6%
–
91 Qwen3-Max-ThinkingAlibaba
83.6%
$0.78 / $3.90
93 GPT-5.4 mini (extra-high reasoning)OpenAI
83.3%
$0.75 / $4.50
93 GPT-5.6 Sol (high reasoning)OpenAI
83.3%
$4 / $20
95 GPT-5.1-Codex (high reasoning)OpenAI
83.0%
$1.25 / $10
96 GPT-5.1 (high reasoning)OpenAI
81.9%
$1.25 / $10
97 Qwen3.5-2B (no reasoning)Alibaba
81.6%
–
98 Kimi K2.5 (no reasoning)Moonshot AI
81.3%
$0.57 / $2.85
99 GPT-5.6 Sol (medium reasoning)OpenAI
81.0%
$4 / $20
100 o3OpenAI
80.7%
$2 / $8
101 Gemini 3 Flash Preview (reasoning on)Google
80.4%
$0.50 / $3
101 GPT-5.6 Terra (extra-high reasoning)OpenAI
80.4%
$2 / $12
103 Qwen3-Coder-NextAlibaba
79.5%
$0.18 / $0.90
103 Claude Sonnet 4.6 (no reasoning)Anthropic
79.5%
$3 / $15
105 DeepSeek-V3.2 (no reasoning)DeepSeek
78.9%
$0.30 / $0.96
106 GPT-5.6 Terra (high reasoning)OpenAI
78.4%
$2 / $12
107 Claude Sonnet 4.5 (reasoning on)Anthropic
78.1%
$3 / $15
108 GLM-4.6 (no reasoning)Z.ai
76.9%
$0.50 / $2
109 GPT-5.4 nano (extra-high reasoning)OpenAI
76.0%
$0.20 / $1.25
109 GPT-5.6 Sol (low reasoning)OpenAI
76.0%
$4 / $20
111 Nova 2 Lite (medium reasoning)Amazon
75.7%
$0.30 / $2.50
111 Claude Sonnet 4.6 (max reasoning)Anthropic
75.7%
$3 / $15
111 Grok Code Fast 1xAI
75.7%
–
114 Grok 4xAI
74.9%
–
115 GPT-5.4 (low reasoning)OpenAI
74.6%
$2.50 / $15
116 Qwen3-MaxAlibaba
74.3%
$0.78 / $3.90
116 GPT-5.2 (medium reasoning)OpenAI
74.3%
$1.75 / $14
118 Claude Opus 4.7 (no reasoning)Anthropic
74.0%
$5 / $25
119 Kimi K2 (0905)Moonshot AI
73.4%
$0.60 / $2.50
120 Nova 2 Lite (high reasoning)Amazon
72.8%
$0.30 / $2.50
120 GPT-5.6 Terra (medium reasoning)OpenAI
72.8%
$2 / $12
122 MiMo-V2.5-Pro (no reasoning)Xiaomi
72.5%
$0.43 / $0.87
123 Nova 2 Lite (low reasoning)Amazon
71.9%
$0.30 / $2.50
124 Claude Opus 4.1 (reasoning on)Anthropic
71.4%
$15 / $75
125 GPT-5 mini (medium reasoning)OpenAI
71.1%
$0.25 / $2
126 Mercury 2Inception
70.8%
$0.25 / $0.75
127 Claude Sonnet 4.5 (no reasoning)Anthropic
70.5%
$3 / $15
127 GLM-4.6 (reasoning on)Z.ai
70.5%
$0.50 / $2
129 Grok 4.20 (0309, non-reasoning, no reasoning)xAI
69.6%
–
130 GPT-5.5 (no reasoning)OpenAI
69.3%
$5 / $30
131 Qwen3.5-2B (reasoning on)Alibaba
69.0%
–
132 GPT-5 mini (high reasoning)OpenAI
68.4%
$0.25 / $2
133 Gemini 3 Pro Preview (low reasoning)Google
68.1%
–
134 Nemotron 3 Super 120B A12BNVIDIA
67.8%
$0.085 / $0.40
135 Hy3 Preview (no reasoning)Tencent
67.5%
$0.18 / $0.60
136 GPT-5 (minimal reasoning)OpenAI
67.0%
$1.25 / $10
137 gpt-oss-120b (high reasoning)OpenAI
65.8%
$0.15 / $0.60
137 Grok 4.3 (no reasoning)xAI
65.8%
$1.25 / $2.50
137 Grok 4 Fast (reasoning)xAI
65.8%
–
140 Gemma 4 31B (no reasoning)Google
65.5%
$0.14 / $0.40
141 Qwen3.5-0.8B (no reasoning)Alibaba
65.2%
–
142 Claude Sonnet 4 (reasoning on)Anthropic
64.6%
$3 / $15
143 Grok 4.1 Fast (non-reasoning, no reasoning)xAI
63.7%
–
143 Grok 4 Fast (non-reasoning, no reasoning)xAI
63.7%
–
145 GPT-5.1-Codex-Mini (high reasoning)OpenAI
62.9%
$0.25 / $2
146 o1OpenAI
62.6%
$15 / $60
147 Nova 2 Lite (no reasoning)Amazon
62.0%
$0.30 / $2.50
148 Kimi K2 (0711)Moonshot AI
61.1%
$0.57 / $2.30
149 GPT-5.6 Terra (low reasoning)OpenAI
60.5%
$2 / $12
150 gpt-oss-20b (high reasoning)OpenAI
60.2%
$0.03 / $0.15
151 Gemma 4 31BGoogle
59.9%
$0.14 / $0.40
151 Grok 4.20 (no reasoning)xAI
59.9%
$1.25 / $2.50
153 Gemini 3.5 Flash (minimal reasoning)Google
58.8%
$1.50 / $9
154 o4-mini (high reasoning)OpenAI
55.6%
$1.10 / $4.40
155 Claude Haiku 4.5 (reasoning on)Anthropic
54.7%
$1 / $5
156 Qwen3-VL-235B-A22B-ThinkingAlibaba
54.1%
$0.40 / $4
156 Gemini 2.5 ProGoogle
54.1%
$1.25 / $10
158 Qwen3-235B-A22B-Thinking-2507Alibaba
53.2%
$0.30 / $3
159 GPT-4.1 miniOpenAI
52.9%
$0.40 / $1.60
160 GPT-5.4 nano (medium reasoning)OpenAI
52.6%
$0.20 / $1.25
161 Claude Sonnet 4 (no reasoning)Anthropic
52.3%
$3 / $15
162 Magistral Medium 1.2Mistral AI
52.0%
–
163 gpt-oss-20b (low reasoning)OpenAI
50.3%
$0.03 / $0.15
164 GPT-5.5 Instant (2026-05-26)OpenAI
49.4%
–
165 Qwen3.5-0.8B (reasoning on)Alibaba
47.7%
–
166 DeepSeek-V3-0324DeepSeek
47.1%
$0.25 / $1
166 GPT-4.1OpenAI
47.1%
$2 / $8
168 GPT-5.1 (no reasoning)OpenAI
46.5%
$1.25 / $10
168 GPT-5.2 (no reasoning)OpenAI
46.5%
$1.75 / $14
168 GLM-4.5-AirZ.ai
46.5%
$0.14 / $0.86
171 Qwen3-VL-32B-ThinkingAlibaba
45.6%
–
171 Gemini 2.5 Flash Preview (09-2025, reasoning on)Google
45.6%
–
173 gpt-oss-120b (low reasoning)OpenAI
45.0%
$0.15 / $0.60
174 Qwen3-Coder-480B-A35BAlibaba
43.6%
$0.35 / $1.50
174 Gemma 4 26B A4BGoogle
43.6%
$0.10 / $0.30
176 Gemini 3 Flash Preview (no reasoning)Google
43.3%
$0.50 / $3
177 GLM-4.5Z.ai
43.0%
$0.60 / $2.20
178 Qwen3-Next-80B-A3B-ThinkingAlibaba
41.5%
$0.15 / $1.20
179 Mistral Small 4Mistral AI
41.2%
$0.15 / $0.60
180 Nemotron 3 Nano 30B A3B (reasoning on)NVIDIA
40.9%
$0.05 / $0.20
181 Mistral Medium 3.1Mistral AI
40.6%
$0.40 / $2
182 Gemma 4 26B A4B (no reasoning)Google
40.4%
$0.10 / $0.30
183 DeepSeek-V3.1 (reasoning on)DeepSeek
37.4%
$0.55 / $1.65
184 DeepSeek-V3.1-Terminus (no reasoning)DeepSeek
37.1%
$0.27 / $1
184 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek
37.1%
$0.27 / $1
186 DeepSeek-R1-0528DeepSeek
36.6%
$0.50 / $2.18
186 GPT-5.4 mini (medium reasoning)OpenAI
36.6%
$0.75 / $4.50
186 GPT-5 nano (high reasoning)OpenAI
36.6%
$0.05 / $0.40
189 Gemma 4 12BGoogle
36.3%
–
190 GPT-5.4 (no reasoning)OpenAI
36.0%
$2.50 / $15
191 Qwen3-VL-235B-A22B-InstructAlibaba
35.1%
$0.30 / $1.50
192 DeepSeek-V3.1 (no reasoning)DeepSeek
34.8%
$0.55 / $1.65
192 GPT-5.4 nano (no reasoning)OpenAI
34.8%
$0.20 / $1.25
194 Qwen2.5-72B-InstructAlibaba
34.5%
$0.36 / $0.40
194 Qwen3-14B (reasoning on)Alibaba
34.5%
$0.12 / $0.24
194 Qwen3-Coder-30B-A3B-InstructAlibaba
34.5%
$0.07 / $0.28
197 DeepSeek-V3.2-Exp (no reasoning)DeepSeek
33.9%
$0.27 / $0.41
197 DeepSeek-V3.2-Exp (reasoning on)DeepSeek
33.9%
$0.27 / $0.41
199 Qwen3-235B-A22B-Instruct-2507Alibaba
33.3%
$0.15 / $0.75
200 Mistral Large 2 (2407)Mistral AI
33.0%
$2 / $6
201 Qwen3-Max-PreviewAlibaba
32.7%
–
202 Claude Haiku 4.5 (no reasoning)Anthropic
32.5%
$1 / $5
203 Qwen3-14B (no reasoning)Alibaba
32.2%
$0.12 / $0.24
204 Gemma 4 12B (no reasoning)Google
31.9%
–
204 GPT-5 mini (minimal reasoning)OpenAI
31.9%
$0.25 / $2
206 Gemini 2.5 Flash (reasoning on)Google
31.6%
$0.30 / $2.50
206 GLM-4.6V (reasoning on)Z.ai
31.6%
$0.30 / $0.90
208 Gemini 3.1 Flash-Lite PreviewGoogle
31.3%
$0.25 / $1.50
208 o3-mini (high reasoning)OpenAI
31.3%
$1.10 / $4.40
210 Gemini 2.5 Flash-Lite Preview (09-2025, reasoning on)Google
30.7%
–
210 GLM-4.6V (no reasoning)Z.ai
30.7%
$0.30 / $0.90
212 Gemini 2.5 Flash-Lite Preview (09-2025, no reasoning)Google
30.4%
–
212 GPT-5 nano (medium reasoning)OpenAI
30.4%
$0.05 / $0.40
214 Qwen3-32B (reasoning on)Alibaba
29.8%
$0.14 / $0.40
215 Mistral Small 3.2 24BMistral AI
29.5%
$0.094 / $0.25
216 Qwen3-VL-32B-InstructAlibaba
29.2%
$0.10 / $0.42
216 Qwen3-VL-8B-InstructAlibaba
29.2%
$0.12 / $0.46
218 GPT-4o (2024-08-06)OpenAI
28.9%
$2.50 / $10
219 o3-miniOpenAI
28.7%
$1.10 / $4.40
220 Gemini 2.5 Flash Preview (09-2025, no reasoning)Google
28.4%
–
220 Devstral Small 1.1Mistral AI
28.4%
–
222 Qwen3-30B-A3B-Thinking-2507Alibaba
28.1%
$0.20 / $2.40
223 Qwen3-8B (reasoning on)Alibaba
27.8%
$0.12 / $0.46
223 Magistral Small 1.2Mistral AI
27.8%
–
225 Qwen3-235B-A22B (no reasoning)Alibaba
27.2%
$0.46 / $1.82
225 Ministral 3 14BMistral AI
27.2%
$0.20 / $0.20
227 Qwen3-4B-Instruct-2507Alibaba
26.6%
–
227 Llama 3.3 70B InstructMeta
26.6%
$0.59 / $0.79
227 Ministral 3 8BMistral AI
26.6%
$0.15 / $0.15
230 Qwen3-30B-A3B (reasoning on)Alibaba
26.0%
$0.12 / $0.50
230 Gemma 4 E4B (no reasoning)Google
26.0%
–
232 GPT-5 nano (minimal reasoning)OpenAI
25.7%
$0.05 / $0.40
233 Qwen3-4B-Thinking-2507Alibaba
25.4%
–
233 Nemotron 3 Nano 30B A3B (no reasoning)NVIDIA
25.4%
$0.05 / $0.20
235 Mistral Small 3.1 24BMistral AI
25.1%
$0.35 / $0.56
236 Qwen3-8B (no reasoning)Alibaba
24.9%
$0.12 / $0.46
236 Devstral 2Mistral AI
24.9%
$0.40 / $2
236 Ministral 3 3BMistral AI
24.9%
$0.10 / $0.10
239 Mistral Large 3Mistral AI
24.6%
$0.50 / $1.50
240 Mistral Medium 3Mistral AI
24.3%
$0.40 / $2
241 Qwen3-235B-A22B (reasoning on)Alibaba
24.0%
$0.46 / $1.82
242 Qwen3-VL-4B-InstructAlibaba
23.4%
–
242 Devstral Small 2Mistral AI
23.4%
–
242 GPT-5.4 mini (no reasoning)OpenAI
23.4%
$0.75 / $4.50
245 DeepSeek-V3DeepSeek
22.8%
$0.26 / $1.03
246 Qwen3-VL-8B-ThinkingAlibaba
22.5%
$0.18 / $2.10
246 GLM-4.5V (reasoning on)Z.ai
22.5%
$0.60 / $1.80
248 Qwen3-30B-A3B (no reasoning)Alibaba
22.2%
$0.12 / $0.50
248 Gemma 4 E2B (no reasoning)Google
22.2%
–
250 Qwen3-Next-80B-A3B-InstructAlibaba
21.6%
$0.10 / $1.10
251 Qwen3-Omni-30B-A3B-ThinkingAlibaba
21.3%
–
252 Llama 3.2 3B InstructMeta
21.1%
$0.05 / $0.33
253 Gemma 4 E2BGoogle
20.8%
–
253 Gemma 4 E4BGoogle
20.8%
–
255 Qwen3-VL-30B-A3B-ThinkingAlibaba
19.9%
$0.29 / $1
255 Devstral MediumMistral AI
19.9%
–
257 Mistral Small 3Mistral AI
19.6%
$0.05 / $0.08
257 GLM-4.5V (no reasoning)Z.ai
19.6%
$0.60 / $1.80
259 Qwen3-VL-30B-A3B-InstructAlibaba
19.0%
$0.15 / $0.60
259 Gemini 2.5 Flash-Lite (no reasoning)Google
19.0%
$0.10 / $0.40
261 Gemini 2.5 Flash-Lite (reasoning on)Google
18.4%
$0.10 / $0.40
261 Mistral Small 4 (no reasoning)Mistral AI
18.4%
$0.15 / $0.60
263 Llama 4 MaverickMeta
17.8%
$0.27 / $0.85
264 Nova Lite 1.0Amazon
17.5%
$0.06 / $0.24
265 GPT-4.1 nanoOpenAI
17.3%
$0.10 / $0.40
266 Qwen3-Omni-30B-A3B-InstructAlibaba
16.4%
–
266 Llama 3.1 8B InstructMeta
16.4%
$0.05 / $0.08
268 Qwen3-VL-4B-ThinkingAlibaba
15.5%
–
268 Llama 4 ScoutMeta
15.5%
$0.18 / $0.59
270 Command ACohere
15.2%
$2.50 / $10
270 Llama 3.1 70B InstructMeta
15.2%
$0.40 / $0.40
272 Gemini 2.5 Flash (no reasoning)Google
14.9%
$0.30 / $2.50
273 Nova Micro 1.0Amazon
14.0%
$0.035 / $0.14
273 Nova Pro 1.0Amazon
14.0%
$0.80 / $3.20
275 DeepSeek-R1DeepSeek
11.4%
$0.70 / $2.50
276 Gemma 3 12BGoogle
10.8%
$0.05 / $0.15
277 Gemma 3 27BGoogle
10.5%
$0.12 / $0.20
278 Qwen3-30B-A3B-Instruct-2507Alibaba
10.2%
$0.09 / $0.30
279 Gemma 3 270MGoogle
9.1%
–
280 Gemma 3 4BGoogle
5.0%
$0.05 / $0.10
281 DeepSeek-V3.2-SpecialeDeepSeek
0.0%
–
281 Llama 3.2 1B InstructMeta
0.0%
$0.027 / $0.20
281 Phi-4Microsoft
0.0%
$0.07 / $0.14
281 Kimi Linear 48B A3B InstructMoonshot AI
0.0%
–
281 GPT-5 ChatOpenAI
0.0%
–
281 Reka Flash 3Reka AI
0.0%
$0.10 / $0.20

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Results as published by Artificial Analysis; we do not re-run them.

What it measures

Share of simulated telecom support tasks resolved with tools, as run by Artificial Analysis with its own harness and prompts.

What it does not measure

Not your policies or systems; no longer run on new models.

282 results from Artificial Analysis not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Not on sale through the API providers we track: 268
  • A different snapshot or variant from the model we list: 13
  • An unusual combination of settings: 1

Data sourced from Artificial Analysis. Licence: Artificial Analysis commercial data licence.