Leading or double-barrelled questions: Claude Haiku 4.5
8 results · % of tasks, lower is better · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight
- 1≈ Claude Opus 5.5Anthropic 0.0%
- 1≈ GPT-6.1 SolOpenAI 0.0%
- 3 Claude Haiku 5.5Anthropic 33.3%
- 3 Mistral Large 4Mistral AI 33.3%
- 5 Gemini 3.1 Pro PreviewGoogle 44.4%
- 6 Mistral Medium 3.5Mistral AI 55.6%
- 7 Claude Haiku 4.5Anthropic 66.7%
- 8 Gemini 3.5 Flash-LiteGoogle 77.8%