0.08
Confirmed task success · rank 4 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- 0.06 to 0.09
- Sample
- 21992 observations
- Configuration
- Hy4 preview
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.07
Praise over complaint · rank 8 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- 0.04 to 0.11
- Sample
- 7690 observations
- Configuration
- Hy4 preview
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.01
Steerability · rank 17 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.03 to 0.01
- Sample
- 24074 observations
- Configuration
- Hy4 preview
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.00 to 0.00
- Sample
- 4165961 observations
- Configuration
- Hy4 preview
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.