Benchmarks / Microsoft STATE-Bench
Reported by Microsoft STATE-Bench
Microsoft STATE-Bench
Mean rating of the conversation from the simulated user's side.
- Results dated
- 1 Jul 2026
- Models
- 2
- Unit
- score from 1 to 5
- Licence
- MIT License
| # | Model | STATE-Bench v0.8: user experience score from 1 to 5, higher is better |
|---|---|---|
| 1 | GPT-5.4OpenAI · gpt-5.4 |
3.67
|
| 2 | GPT-5.5OpenAI · gpt-5.5 |
3.62
|
Results as published by Microsoft STATE-Bench; we do not re-run them.
What it measures
Mean rating of the conversation from the simulated user's side.
What it does not measure
Not real customer satisfaction: the ratings come from a model-based judge.
Source
Failures
Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.
Not ranked
- GPT 5.1: memory-track entries are agent-learning systems, not models
- GPT 5.1 + Foundry Memory: memory-track entries are agent-learning systems, not models
- GPT 5.4: memory-track entries are agent-learning systems, not models
- GPT 5.4 + Foundry Memory: memory-track entries are agent-learning systems, not models
STATE-Bench by Microsoft and STATE-Bench contributors, MIT License.