Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

τ²-Bench Telecom benchmark: model scores and methodology

A legacy agentic tool-use benchmark for conversational work in a telecom environment.

Current published leader
GLM-5.2 (max)
Top score
99%4-way tie
Primary metric
Pass rate · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

Published τ²-Bench Telecom model scores in this snapshot

215 configurations · 160 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 215 current configurations with a reported τ²-Bench Telecom score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on τ²-Bench Telecom

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. These results sit in a narrow band, so the scale below is zoomed — read each mark against the labelled axis, not the left edge. Marks are placed on the full source precision, so two models sharing a rounded label can still sit at different points. Score labels reproduce Artificial Analysis's rounded display value.

  1. GLM-5.2 max 99%
  2. JT-35B-Flash 99%
  3. Claude Fable 5 Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 99%
  4. Step 3.7 Flash 99%
  5. Qwen3.6 Plus 98%
  6. DeepSeek V4 Pro Reasoning, Max Effort 96%
  7. DeepSeek V4 Flash Reasoning, High Effort 96%
  8. Gemini 3.1 Pro (Preview) 96%
  9. Gemini 3.5 Flash medium 96%
  10. Qwen3.5 397B A17B Reasoning 96%
  11. Qwen3.6 35B A3B Reasoning 95%
  12. Qwen3.7 Max 95%
  13. MiMo-V2.5-Pro 94%
  14. Mistral Medium 3.5 94%
  15. Qwen3.6 27B Reasoning 94%
  16. Kimi K2.6 Non-reasoning 94%
  17. Qwen3.5 122B A10B Reasoning 94%
  18. MiMo-V2-Flash Feb 2026 93%
  19. JT-MINI 93%
  20. Qwen3.7 Plus 93%
90.8 95.3 99.9

Pass rate · higher is better · zoomed scale

Current configurations with a reported τ²-Bench Telecom value; missing values are omitted.

How to interpret the result

What do τ²-Bench Telecom results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher pass rate is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read τ²-Bench Telecom as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationPass rate
1 GLM-5.2 (max)Leader max 99%
2 JT-35B-Flash Published configuration 99%
3 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 99%
4 Step 3.7 Flash Published configuration 99%
5 Qwen3.6 Plus Published configuration 98%
6 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 96%
7 DeepSeek V4 Flash (Reasoning, High Effort) Reasoning, High Effort 96%
8 Gemini 3.1 Pro (Preview) Published configuration 96%
9 Gemini 3.5 Flash (medium) medium 96%
10 Qwen3.5 397B A17B (Reasoning) Reasoning 96%
11 Gemini 3.5 Flash (high) high 95%
12 Qwen3.6 35B A3B (Reasoning) Reasoning 95%
13 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort 95%
14 Qwen3.7 Max Published configuration 95%
15 DeepSeek V4 Flash (Non-reasoning) Non-reasoning 94%
16 DeepSeek V4 Pro (Reasoning, High Effort) Reasoning, High Effort 94%
17 MiMo-V2.5-Pro Published configuration 94%
18 Mistral Medium 3.5 Published configuration 94%
19 Qwen3.6 27B (Reasoning) Reasoning 94%
20 Kimi K2.6 (Non-reasoning) Non-reasoning 94%
21 Qwen3.5 122B A10B (Reasoning) Reasoning 94%
22 Qwen3.6 27B (Non-reasoning) Non-reasoning 94%
23 MiMo-V2-Flash (Feb 2026) Feb 2026 93%
24 JT-MINI Published configuration 93%
25 Qwen3.7 Plus Published configuration 93%
Show the remaining 190 configurations
#ModelConfigurationPass rate
26 Hy3-preview (Reasoning) Reasoning 93%
27 Nova 2.0 Pro Preview (medium) medium 93%
28 Ring-2.6-1T Published configuration 92%
29 Qwen3.5 4B (Reasoning) Reasoning 92%
30 Muse Spark Published configuration 92%
31 DeepSeek V4 Pro (Non-reasoning) Non-reasoning 91%
32 Grok 4.3 (medium) medium 91%
33 MiMo-V2-Omni Published configuration 91%
34 MiMo-V2.5 Published configuration 91%
35 Nova 2.0 Pro Preview (low) low 91%
36 Kimi K2.7 Code Published configuration 90%
37 Trinity Large Thinking Published configuration 90%
38 Ling-2.6-1T Published configuration 90%
39 KAT Coder Pro V2 Published configuration 89%
40 Grok 4.3 (low) low 89%
41 MiniMax-M3 Published configuration 89%
42 KAT-Coder-Pro V1 Published configuration 89%
43 Qwen3.5 Omni Plus Published configuration 88%
44 MiMo-V2-Omni-0327 Published configuration 88%
45 MiniCPM-V 4.6 1.3B Published configuration 88%
46 Qwen3.5 4B (Non-reasoning) Non-reasoning 88%
47 HyperCLOVA X SEED Think (32B) 32B 87%
48 Qwen3.5 9B (Reasoning) Reasoning 87%
49 Mi:dm K 2.5 Pro Published configuration 87%
50 GPT-5.6 Terra max 86%
51 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning 86%
52 Solar Pro 3 Published configuration 86%
53 GPT-5.3 Codex (xhigh) xhigh 86%
54 Ling 2.6 Flash Published configuration 86%
55 GPT-5.6 Sol max 85%
56 Qwen3.5 9B (Non-reasoning) Non-reasoning 85%
57 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning 85%
58 GPT-5.6 Sol (xhigh) xhigh 85%
59 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning 85%
60 Qwen3.5 Omni Flash Published configuration 85%
61 MiMo-V2-Flash (Non-reasoning) Non-reasoning 84%
62 Qwen3.5 397B A17B (Non-reasoning) Non-reasoning 84%
63 ERNIE 5.0 Thinking Preview Published configuration 84%
64 GPT-5.6 Sol (high) high 83%
65 Nemotron 3 Ultra 550B A55B (Reasoning) Reasoning 83%
66 MiniCPM5-1B (Non-reasoning) Non-reasoning 82%
67 Nex-N2-Pro Published configuration 82%
68 Qwen3.5 2B (Non-reasoning) Non-reasoning 82%
69 GPT-5.6 Sol (medium) medium 81%
70 MiniCPM5-1B (Reasoning) Reasoning 81%
71 Tri-21B-Think Published configuration 81%
72 Command A+ Published configuration 81%
73 o3 Published configuration 81%
74 GPT-5.6 Terra (xhigh) xhigh 80%
75 Nova 2.0 Omni (medium) medium 80%
76 LongCat Flash Lite Published configuration 80%
77 Qwen3 Coder Next Published configuration 80%
78 Claude Sonnet 4.6 (Non-reasoning, Low Effort) Non-reasoning, Low Effort 79%
79 GPT-5.6 Terra (high) high 78%
80 EXAONE 4.5 33B Published configuration 78%
81 GPT-5.6 Sol (low) low 76%
82 Nova 2.0 Lite (medium) medium 76%
83 K-EXAONE (Reasoning) Reasoning 74%
84 GPT-5.6 Terra (medium) medium 73%
85 Nova 2.0 Lite (high) high 73%
86 MiMo-V2.5-Pro (Non-reasoning) Non-reasoning 73%
87 Nova 2.0 Lite (low) low 72%
88 Nova 2.0 Pro Preview (Non-reasoning) Non-reasoning 72%
89 Mercury 2 Published configuration 71%
90 Apriel-v1.6-15B-Thinker Published configuration 69%
91 Qwen3.5 2B (Reasoning) Reasoning 69%
92 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 68%
93 Nova 2.0 Omni (low) low 68%
94 Hy3-preview (Non-reasoning) Non-reasoning 68%
95 Grok 4.3 (Non-reasoning) Non-reasoning 66%
96 gpt-oss-120b (high) high 66%
97 Gemma 4 31B (Non-reasoning) Non-reasoning 65%
98 Qwen3.5 0.8B (Non-reasoning) Non-reasoning 65%
99 HyperNova 60B 2605 Published configuration 63%
100 Nova 2.0 Lite (Non-reasoning) Non-reasoning 62%
101 GPT-5.6 Terra (low) low 61%
102 gpt-oss-20b (high) high 60%
103 Gemma 4 31B (Reasoning) Reasoning 60%
104 K-EXAONE (Non-reasoning) Non-reasoning 59%
105 Gemini 3.5 Flash (minimal) minimal 59%
106 Doubao Seed Code Published configuration 58%
107 Claude 4.5 Haiku (Reasoning) Reasoning 55%
108 Gemini 2.5 Pro Published configuration 54%
109 Nemotron Cascade 2 30B A3B Published configuration 53%
110 Magistral Medium 1.2 Published configuration 52%
111 gpt-oss-20b (low) low 50%
112 Mi:dm K 2.5 Pro Preview Published configuration 49%
113 Solar Open 100B (Reasoning) Reasoning 48%
114 Qwen3.5 0.8B (Reasoning) Reasoning 48%
115 Sarvam 105B (high) high 47%
116 Motif-2-12.7B-Reasoning Published configuration 46%
117 Nemotron 3 Nano Omni 30B A3B Reasoning Published configuration 45%
118 gpt-oss-120b (low) low 45%
119 Nova 2.0 Omni (Non-reasoning) Non-reasoning 45%
120 Gemma 4 26B A4B (Reasoning) Reasoning 44%
121 Granite 4.1 30B Published configuration 42%
122 Qwen3 Next 80B A3B (Reasoning) Reasoning 42%
123 Mistral Small 4 (Reasoning) Reasoning 41%
124 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) Reasoning 41%
125 Gemma 4 26B A4B (Non-reasoning) Non-reasoning 40%
126 Nova Premier Published configuration 38%
127 North Mini Code Published configuration 37%
128 Gemma 4 12B (Reasoning) Reasoning 36%
129 Sarvam 30B (high) high 35%
130 Claude 4.5 Haiku (Non-reasoning) Non-reasoning 32%
131 Gemma 4 12B (Non-reasoning) Non-reasoning 32%
132 Gemini 3.1 Flash-Lite Published configuration 31%
133 Llama Nemotron Super 49B v1.5 (Reasoning) Reasoning 28%
134 NVIDIA Nemotron 3 Nano 4B Published configuration 28%
135 Falcon-H1R-7B Published configuration 28%
136 Granite 4.1 8B Published configuration 28%
137 K2-V2 (high) high 28%
138 Magistral Small 1.2 Published configuration 28%
139 Ministral 3 14B Published configuration 27%
140 Hermes 4 - Llama-3.1 405B (Non-reasoning) Non-reasoning 27%
141 INTELLECT-3 Published configuration 27%
142 Llama 3.3 Instruct 70B Published configuration 27%
143 Ministral 3 8B Published configuration 27%
144 Gemma 4 E4B (Non-reasoning) Non-reasoning 26%
145 K2 Think V2 Published configuration 25%
146 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) Non-reasoning 25%
147 Llama Nemotron Super 49B v1.5 (Non-reasoning) Non-reasoning 25%
148 Devstral 2 Published configuration 25%
149 K2-V2 (medium) medium 25%
150 Ministral 3 3B Published configuration 25%
151 Mistral Large 3 Published configuration 25%
152 Devstral Small 2 Published configuration 23%
153 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) Non-reasoning 23%
154 Llama 3.1 Nemotron Instruct 70B Published configuration 23%
155 Granite 4.0 1B Published configuration 23%
156 Hermes 4 - Llama-3.1 70B (Reasoning) Reasoning 23%
157 Gemma 4 E2B (Non-reasoning) Non-reasoning 22%
158 Hermes 4 - Llama-3.1 405B (Reasoning) Reasoning 22%
159 NVIDIA Nemotron Nano 9B V2 (Reasoning) Reasoning 22%
160 Hermes 4 - Llama-3.1 70B (Non-reasoning) Non-reasoning 22%
161 Nanbeige4.1-3B Published configuration 22%
162 Qwen3 Next 80B A3B Instruct Published configuration 22%
163 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) Reasoning 21%
164 Olmo 3.1 32B Instruct Published configuration 21%
165 Qwen3 Omni 30B A3B (Reasoning) Reasoning 21%
166 Gemma 4 E2B (Reasoning) Reasoning 21%
167 Gemma 4 E4B (Reasoning) Reasoning 21%
168 K2-V2 (low) low 21%
169 Exaone 4.0 1.2B (Non-reasoning) Non-reasoning 20%
170 Granite 4.0 H 1B Published configuration 20%
171 Granite 4.1 3B Published configuration 20%
172 LFM2.5-1.2B-Thinking Published configuration 20%
173 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) Non-reasoning 19%
174 Llama 3.1 Instruct 405B Published configuration 19%
175 Mistral Small 4 (Non-reasoning) Non-reasoning 18%
176 Llama 4 Maverick Published configuration 18%
177 EXAONE 4.0 32B (Reasoning) Reasoning 17%
178 Granite 4.0 H Small Published configuration 17%
179 Exaone 4.0 1.2B (Reasoning) Reasoning 16%
180 Qwen3 Omni 30B A3B Instruct Published configuration 16%
181 LFM2.5-8B-A1B Published configuration 16%
182 Step3 VL 10B Published configuration 16%
183 Jamba Reasoning 3B Published configuration 16%
184 Llama 4 Scout Published configuration 15%
185 Command A Published configuration 15%
186 Granite 4.0 H 350M Published configuration 15%
187 Llama 3.2 Instruct 11B (Vision) Vision 15%
188 Nova Micro Published configuration 14%
189 Jamba 1.7 Large Published configuration 13%
190 LFM2 2.6B Published configuration 13%
191 Granite 4.0 350M Published configuration 13%
192 Ling-mini-2.0 Published configuration 13%
193 Apertus 70B Instruct Published configuration 13%
194 Granite 4.0 Micro Published configuration 13%
195 Jamba 1.7 Mini Published configuration 13%
196 Olmo 3 7B Instruct Published configuration 13%
197 Apertus 8B Instruct Published configuration 11%
198 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) Reasoning 11%
199 LFM2 24B A2B Published configuration 11%
200 LFM2.5-1.2B-Instruct Published configuration 11%
201 LFM2 8B A1B Published configuration 11%
202 Gemma 3 270M Published configuration 9%
203 LFM2.5-VL-1.6B Published configuration 8%
204 Phi-4 Mini Instruct Published configuration 8%
205 EXAONE 4.0 32B (Non-reasoning) Non-reasoning 4%
206 ERNIE 4.5 300B A47B Published configuration 0%
207 Kimi Linear 48B A3B Instruct Published configuration 0%
208 Molmo 7B-D Published configuration 0%
209 Molmo2-8B Published configuration 0%
210 Olmo 3 7B Think Published configuration 0%
211 Olmo 3.1 32B Think Published configuration 0%
212 Phi-4 Published configuration 0%
213 Reka Flash 3 Published configuration 0%
214 Ring-flash-2.0 Published configuration 0%
215 Tiny Aya Global Published configuration 0%

Showing the top 25 of 215 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

A legacy agentic tool-use benchmark for conversational work in a telecom environment. Artificial Analysis replaced this constituent with τ³-Banking in Intelligence Index v4.1; do not merge the two tracks.

How to read the score

The published metric is Pass rate. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same tau2 field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official τ²-Bench Telecom resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗