Benchmarks / tau2-bench
Reported by tau2-bench
tau2-bench
Share of banking questions answered from a policy knowledge base tasks completed within policy in all four trials.
- Results dated
- 5 May 2026 to 3 Aug 2026
- Models
- 21
- Unit
- % of tasks
- Licence
- MIT License
| # | Model | tau2 v1.0.1, banking knowledge, alltools retrieval: consistency (pass^4) % of tasks, higher is better |
|---|---|---|
| 1 | Qwen3.8 Max (0902)Alibaba · qwen3.8-max-0902 |
35.1%
|
| 2 | Claude Opus 5Anthropic · claude-opus-5 |
32.0%
|
| 2 | Grok 4.5xAI · grok-4.5 |
32.0%
|
| 4 | GPT-5.5OpenAI · gpt-5.5 |
29.9%
|
| 5 | Claude Fable 5Anthropic · claude-fable-5 |
28.9%
|
| 6 | GPT-5.6 SolOpenAI · gpt-5.6-sol |
27.8%
|
| 7 | Claude Opus 4.7Anthropic · claude-opus-4.7 |
24.7%
|
| 8 | Claude Opus 4.8Anthropic · claude-opus-4.8 |
22.7%
|
| 9 | GPT-5.4OpenAI · gpt-5.4 |
21.6%
|
| 10 | Muse Spark 1.1Meta · muse-spark-1.1 |
20.6%
|
| 11 | GPT-5.2OpenAI · gpt-5.2 |
18.6%
|
| 12 | Kimi K3Moonshot AI · kimi-k3 |
17.5%
|
| 13 | GLM 5.2Z.ai · glm-5.2 |
13.4%
|
| 14 | InklingThinkingmachines · inkling |
11.3%
|
| 15 | Claude Opus 4.5Anthropic · claude-opus-4.5 |
11.3%
|
| 15 | Claude Opus 4.6Anthropic · claude-opus-4.6 |
11.3%
|
| 17 | Gemini 3.1 Pro PreviewGoogle · gemini-3.1-pro-preview |
9.3%
|
| 18 | Grok 4.20xAI · grok-4.20 |
8.2%
|
| 19 | grok-4-1-fastxAI |
5.2%
|
| 20 | grok-4-fastxAI |
4.1%
|
| 21 | Gemini 2.5 ProGoogle · gemini-2.5-pro |
1.0%
|
Results as published by tau2-bench; we do not re-run them.
What it measures
Share of banking questions answered from a policy knowledge base tasks completed within policy in all four trials.
What it does not measure
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Source
Failures
Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.
Not ranked
- Distyl ButtonAgent: custom agent scaffold, not the standard harness
- RAFT-30B-A3B: custom agent scaffold, not the standard harness
tau2-bench by Sierra Research and contributors, MIT License.