Benchmarks / Microsoft STATE-Bench
Reported by Microsoft STATE-Bench
Microsoft STATE-Bench
Share of customer support tasks completed correctly, averaged over repeated runs.
- Results dated
- 25 May 2026 to 29 May 2026
- Models
- 5
- Unit
- % of tasks
- Licence
- MIT License
| # | Model | STATE-Bench v0.7: customer support task success (pass@1) % of tasks, higher is better |
|---|---|---|
| 1 | GPT-5.4OpenAI · gpt-5.4 |
57.6%
|
| 2 | Claude Opus 4.7Anthropic · claude-opus-4.7 |
51.0%
|
| 3 | GPT-5.4OpenAI · gpt-5.4 |
47.2%
|
| 4 | DeepSeek V4 Pro 0423DeepSeek · deepseek-v4-pro |
45.6%
|
| 5 | Kimi K2.6Moonshot AI · kimi-k2.6 |
45.1%
|
Results as published by Microsoft STATE-Bench; we do not re-run them.
What it measures
Share of customer support tasks completed correctly, averaged over repeated runs.
What it does not measure
Not the model alone: it runs inside STATE-Bench's agent loop against simulated users, and versions of the benchmark are not comparable.
Source
Failures
Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.
Not ranked
- GPT 5.1: memory-track entries are agent-learning systems, not models
- GPT 5.1 + Foundry Memory: memory-track entries are agent-learning systems, not models
- GPT 5.4: memory-track entries are agent-learning systems, not models
- GPT 5.4 + Foundry Memory: memory-track entries are agent-learning systems, not models
STATE-Bench by Microsoft and STATE-Bench contributors, MIT License.