Benchmarks / tau2-bench

Reported by tau2-bench

tau2-bench

Share of banking questions answered from a policy knowledge base tasks completed within policy in all four trials.

Results dated
5 May 2026 to 3 Aug 2026
Models
21
Unit
% of tasks
Licence
MIT License
tau2 v1.0.1, banking knowledge, alltools retrieval: consistency (pass^4), % of tasks, higher is better
#Modeltau2 v1.0.1, banking knowledge, alltools retrieval: consistency (pass^4)
% of tasks, higher is better
1 Qwen3.8 Max (0902)Alibaba · qwen3.8-max-0902
35.1%
2 Claude Opus 5Anthropic · claude-opus-5
32.0%
2 Grok 4.5xAI · grok-4.5
32.0%
4 GPT-5.5OpenAI · gpt-5.5
29.9%
5 Claude Fable 5Anthropic · claude-fable-5
28.9%
6 GPT-5.6 SolOpenAI · gpt-5.6-sol
27.8%
7 Claude Opus 4.7Anthropic · claude-opus-4.7
24.7%
8 Claude Opus 4.8Anthropic · claude-opus-4.8
22.7%
9 GPT-5.4OpenAI · gpt-5.4
21.6%
10 Muse Spark 1.1Meta · muse-spark-1.1
20.6%
11 GPT-5.2OpenAI · gpt-5.2
18.6%
12 Kimi K3Moonshot AI · kimi-k3
17.5%
13 GLM 5.2Z.ai · glm-5.2
13.4%
14 InklingThinkingmachines · inkling
11.3%
15 Claude Opus 4.5Anthropic · claude-opus-4.5
11.3%
15 Claude Opus 4.6Anthropic · claude-opus-4.6
11.3%
17 Gemini 3.1 Pro PreviewGoogle · gemini-3.1-pro-preview
9.3%
18 Grok 4.20xAI · grok-4.20
8.2%
19 grok-4-1-fastxAI
5.2%
20 grok-4-fastxAI
4.1%
21 Gemini 2.5 ProGoogle · gemini-2.5-pro
1.0%

Results as published by tau2-bench; we do not re-run them.

What it measures

Share of banking questions answered from a policy knowledge base tasks completed within policy in all four trials.

What it does not measure

Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.

Failures

Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.

Not ranked

  • Distyl ButtonAgent: custom agent scaffold, not the standard harness
  • RAFT-30B-A3B: custom agent scaffold, not the standard harness

tau2-bench by Sierra Research and contributors, MIT License.