Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

τ³-Banking benchmark: model scores and methodology

An agentic banking customer-support benchmark for policy and knowledge retrieval, reasoning, dialogue, and chained tool calls.

Current published leader
Kimi K3 · Published configuration
Top score
33%5-way tie
Primary metric
Pass rate · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current τ³-Banking model scores

130 configurations · 96 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 130 current configurations with a reported τ³-Banking score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on τ³-Banking

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. Score labels reproduce Artificial Analysis's rounded display value.

  1. Kimi K3 33%
  2. GPT-5.6 Sol 33%
  3. Claude Opus 5 Adaptive Reasoning, High Effort 33%
  4. Grok 4.5 33%
  5. GPT-5.6 Terra 32%
  6. Motif 3 Beta 29%
  7. Claude Sonnet 5 Adaptive Reasoning, Max Effort 28%
  8. JT-4.1 Flash 236B A21B 28%
  9. GPT-5.6 Luna 27%
  10. Claude Fable 5 Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 27%
  11. GLM-5.2 max 27%
  12. DeepSeek V4 Pro Reasoning, Max Effort 26%
  13. Gemini 3.5 Flash high 25%
  14. Muse Spark 1.1 xhigh 25%
  15. Gemini 3.6 Flash high 25%
  16. Inkling xhigh 24%
  17. DeepSeek V4 Flash Reasoning, Max Effort 23%
  18. Hy3 21%
  19. Muse Spark 20%
  20. Kimi K2.7 Code 18%
0.0 16.7 33.4

Pass rate · higher is better

Current configurations with a reported τ³-Banking value; missing values are omitted.

How to interpret the result

What do τ³-Banking results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher pass rate is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read τ³-Banking as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationPass rate
1 Kimi K3Leader Published configuration 33%
2 GPT-5.6 Sol max 33%
3 Claude Opus 5 (Adaptive Reasoning, High Effort) Adaptive Reasoning, High Effort 33%
4 GPT-5.6 Sol (xhigh) xhigh 33%
5 Grok 4.5 high 33%
6 GPT-5.6 Terra max 32%
7 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Adaptive Reasoning, Xhigh Effort 32%
8 GPT-5.6 Sol (high) high 31%
9 Claude Opus 5 Adaptive Reasoning, Max Effort 30%
10 Motif 3 (Beta) Beta 29%
11 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Adaptive Reasoning, Medium Effort 29%
12 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Adaptive Reasoning, Max Effort 28%
13 JT-4.1 Flash 236B A21B Published configuration 28%
14 GPT-5.6 Luna max 27%
15 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 27%
16 GLM-5.2 (max) max 27%
17 GPT-5.6 Sol (medium) medium 26%
18 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 26%
19 Gemini 3.5 Flash (high) high 25%
20 Muse Spark 1.1 (xhigh) xhigh 25%
21 Gemini 3.6 Flash (high) high 25%
22 GPT-5.6 Sol (low) low 24%
23 GPT-5.6 Luna (xhigh) xhigh 24%
24 GPT-5.6 Terra (xhigh) xhigh 24%
25 DeepSeek V4 Pro (Reasoning, High Effort) Reasoning, High Effort 24%
Show the remaining 105 configurations
#ModelConfigurationPass rate
26 Inkling (xhigh) xhigh 24%
27 Claude Opus 5 (Adaptive Reasoning, Low Effort) Adaptive Reasoning, Low Effort 23%
28 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort 23%
29 GPT-5.6 Luna (high) high 22%
30 GPT-5.6 Terra (high) high 22%
31 Hy3 Published configuration 21%
32 DeepSeek V4 Flash (Reasoning, High Effort) Reasoning, High Effort 20%
33 Muse Spark Published configuration 20%
34 GPT-5.6 Terra (medium) medium 19%
35 Kimi K2.7 Code Published configuration 18%
36 Qwen3.7 Plus Published configuration 18%
37 Nex-N2-Pro Published configuration 17%
38 Gemini 3.1 Pro (Preview) Published configuration 16%
39 Gemini 3.5 Flash-Lite Published configuration 16%
40 Qwen3.6 Plus Published configuration 16%
41 GPT-5.6 Sol (Non-reasoning) Non-reasoning 16%
42 GPT-5.6 Terra (low) low 16%
43 GLM-5.2 (Non-reasoning) Non-reasoning 16%
44 GPT-5.6 Luna (medium) medium 15%
45 Qwen3.6 27B (Reasoning) Reasoning 15%
46 Gemma 4 31B (Reasoning) Reasoning 15%
47 Mistral Medium 3.5 Published configuration 14%
48 K-EXAONE (Reasoning) Reasoning 14%
49 Ring-2.6-1T Published configuration 14%
50 Claude Sonnet 5 (Non-reasoning, High Effort) Non-reasoning, High Effort 14%
51 Nemotron 3 Ultra 550B A55B (Reasoning) Reasoning 14%
52 Qwen3.5 122B A10B (Reasoning) Reasoning 14%
53 GPT-5.6 Terra (Non-reasoning) Non-reasoning 13%
54 Qwen3.5 397B A17B (Reasoning) Reasoning 13%
55 MiniMax-M3 Published configuration 13%
56 LongCat 2.0 Published configuration 13%
57 gpt-oss-120b (high) high 12%
58 GPT-5.6 Luna (low) low 12%
59 Gemma 4 26B A4B (Reasoning) Reasoning 12%
60 Agnes 2.5 Pro Alpha Published configuration 12%
61 GPT-5.5 Instant (June 2026) June 2026 12%
62 Step 3.7 Flash Published configuration 11%
63 EXAONE 4.5 33B Published configuration 11%
64 Qwen3.7 Max Published configuration 11%
65 Devstral 2 Published configuration 10%
66 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 10%
67 Nemotron Cascade 2 30B A3B Published configuration 10%
68 Mercury 2 Published configuration 10%
69 Devstral Small 2 Published configuration 10%
70 Gemini 2.5 Pro Published configuration 9%
71 Claude 4.5 Haiku (Reasoning) Reasoning 9%
72 GPT-5.6 Luna (Non-reasoning) Non-reasoning 9%
73 Nova 2.0 Pro Preview (low) low 9%
74 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning 9%
75 Gemini 3.1 Flash-Lite Published configuration 9%
76 Gemma 4 12B (Reasoning) Reasoning 9%
77 MiMo-V2.5-Pro Published configuration 9%
78 Qwen3.6 35B A3B (Reasoning) Reasoning 9%
79 Gemma 4 31B (Non-reasoning) Non-reasoning 8%
80 Qwen3.5 4B (Reasoning) Reasoning 8%
81 Qwen3.5 9B (Reasoning) Reasoning 8%
82 Qwen3.6 27B (Non-reasoning) Non-reasoning 8%
83 Solar Pro 3 Published configuration 8%
84 Grok 4.3 (Non-reasoning) Non-reasoning 8%
85 Nova 2.0 Pro Preview (medium) medium 8%
86 NVIDIA Nemotron 3 Nano 4B Published configuration 7%
87 gpt-oss-20b (high) high 7%
88 Nova 2.0 Lite (high) high 7%
89 Nova 2.0 Pro Preview (Non-reasoning) Non-reasoning 7%
90 DiffusionGemma 26B A4B Published configuration 7%
91 MiMo-V2.5 Published configuration 7%
92 Ministral 3 14B Published configuration 7%
93 North Mini Code Published configuration 6%
94 Qwen3 Next 80B A3B (Reasoning) Reasoning 6%
95 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) Reasoning 6%
96 Command A+ Published configuration 6%
97 G9v3-3B Published configuration 6%
98 KAT Coder Pro V2 Published configuration 6%
99 Mistral Large 3 Published configuration 6%
100 Magistral Medium 1.2 Published configuration 6%
101 Trinity Large Thinking Published configuration 6%
102 Gemma 4 E4B (Reasoning) Reasoning 5%
103 HyperNova 60B 2605 Published configuration 5%
104 K2 Think V2 Published configuration 5%
105 Mistral Small 4 (Reasoning) Reasoning 5%
106 Qwen3 Coder Next Published configuration 5%
107 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning 5%
108 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning 5%
109 Ministral 3 3B Published configuration 5%
110 Gemma 4 E2B (Reasoning) Reasoning 5%
111 Magistral Small 1.2 Published configuration 4%
112 MiniCPM-V 4.6 1.3B Published configuration 4%
113 Qwen3.5 9B (Non-reasoning) Non-reasoning 4%
114 Granite 4.1 30B Published configuration 4%
115 Llama 4 Maverick Published configuration 4%
116 Ministral 3 8B Published configuration 4%
117 Llama 4 Scout Published configuration 3%
118 Granite 4.1 8B Published configuration 3%
119 MiMo-V2-Flash (Non-reasoning) Non-reasoning 3%
120 Qwen3.5 4B (Non-reasoning) Non-reasoning 3%
121 Ling 2.6 Flash Published configuration 3%
122 gpt-oss-120b (low) low 3%
123 Qwen3.5 2B (Reasoning) Reasoning 2%
124 Qwen3.5 2B (Non-reasoning) Non-reasoning 2%
125 Granite 4.1 3B Published configuration 1%
126 Qwen3.5 0.8B (Non-reasoning) Non-reasoning 1%
127 Llama 3.3 Instruct 70B Published configuration 1%
128 Phi-4 Mini Instruct Published configuration 1%
129 Nanbeige4.1-3B Published configuration 0%
130 Qwen3.5 0.8B (Reasoning) Reasoning 0%

Showing the top 25 of 130 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

An agentic banking customer-support benchmark for policy and knowledge retrieval, reasoning, dialogue, and chained tool calls.

How to read the score

The published metric is Pass rate. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same tauBanking field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official τ³-Banking resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗