Benchmarks / tau2-bench

Reported by tau2-bench

tau2-bench

Share of telecom account and technical support tasks completed within policy, averaged over trials.

Results dated
24 Feb 2026 to 26 May 2026
Models
8
Unit
% of tasks
Licence
MIT License
tau2 v1.0.1, telecom: task success (pass^1), % of tasks, higher is better
#Modeltau2 v1.0.1, telecom: task success (pass^1)
% of tasks, higher is better
1 Qwen3.5 397B A17BAlibaba · qwen3.5-397b-a17b
97.8%
2 Claude Opus 4.5Anthropic · claude-opus-4.5
92.3%
3 gemini-3-flashGoogle
91.2%
4 gemini-3-proGoogle
91.0%
5 GPT-5.2OpenAI · gpt-5.2
89.7%
6 GLM 5Z.ai · glm-5
86.8%
7 Claude Sonnet 4.5Anthropic · claude-sonnet-4.5
84.9%
8 GPT-5.2OpenAI · gpt-5.2
57.2%

Results as published by tau2-bench; we do not re-run them.

What it measures

Share of telecom account and technical support tasks completed within policy, averaged over trials.

What it does not measure

Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.

Failures

Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.

Not ranked

  • Distyl ButtonAgent: custom agent scaffold, not the standard harness
  • RAFT-30B-A3B: custom agent scaffold, not the standard harness

tau2-bench by Sierra Research and contributors, MIT License.