Benchmarks / Microsoft STATE-Bench

Reported by Microsoft STATE-Bench

Microsoft STATE-Bench

Mean rating of the conversation from the simulated user's side.

Results dated
1 Jul 2026
Models
2
Unit
score from 1 to 5
Licence
MIT License
STATE-Bench v0.8: user experience, score from 1 to 5, higher is better
#ModelSTATE-Bench v0.8: user experience
score from 1 to 5, higher is better
1 GPT-5.4OpenAI · gpt-5.4
3.67
2 GPT-5.5OpenAI · gpt-5.5
3.62

Results as published by Microsoft STATE-Bench; we do not re-run them.

What it measures

Mean rating of the conversation from the simulated user's side.

What it does not measure

Not real customer satisfaction: the ratings come from a model-based judge.

Failures

Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.

Not ranked

  • GPT 5.1: memory-track entries are agent-learning systems, not models
  • GPT 5.1 + Foundry Memory: memory-track entries are agent-learning systems, not models
  • GPT 5.4: memory-track entries are agent-learning systems, not models
  • GPT 5.4 + Foundry Memory: memory-track entries are agent-learning systems, not models

STATE-Bench by Microsoft and STATE-Bench contributors, MIT License.