-0.09
Confirmed task success · rank 34 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.13 to -0.06
- Sample
- 14957 observations
- Configuration
- Qwen3.7 Plus
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.13
Praise over complaint · rank 35 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.16 to -0.09
- Sample
- 5948 observations
- Configuration
- Qwen3.7 Plus
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.07
Steerability · rank 34 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.10 to -0.05
- Sample
- 22551 observations
- Configuration
- Qwen3.7 Plus
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.01
Tool grounding · rank 38 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.01 to -0.01
- Sample
- 957879 observations
- Configuration
- Qwen3.7 Plus
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,453
Overall · rank 18 of 44
- Unit
- Arena rating, higher is better
- Range
- 1,443 to 1,462
- Sample
- 4591 votes
- Configuration
- Qwen3.7 Plus
- Measured
- 13 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,453
Overall · rank 45 of 177
- Unit
- Arena rating, higher is better
- Range
- 1,449 to 1,456
- Sample
- 41988 votes
- Configuration
- Qwen3.7 Plus
- Measured
- 25 Sep 2026
- Not shown
- Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,445
Business, management and finance · rank 59 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,437 to 1,452
- Sample
- 8423 votes
- Configuration
- Qwen3.7 Plus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,432
Creative writing · rank 40 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,424 to 1,439
- Sample
- 7776 votes
- Configuration
- Qwen3.7 Plus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,485
Expert prompts · rank 32 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,476 to 1,494
- Sample
- 4696 votes
- Configuration
- Qwen3.7 Plus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,447
Instruction following · rank 47 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,441 to 1,453
- Sample
- 15047 votes
- Configuration
- Qwen3.7 Plus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,456
Overall · rank 52 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,451 to 1,460
- Sample
- 42031 votes
- Configuration
- Qwen3.7 Plus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,437
Writing, literature and language · rank 48 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,430 to 1,443
- Sample
- 10803 votes
- Configuration
- Qwen3.7 Plus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.