0.07
Confirmed task success · rank 4 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- 0.04 to 0.09
- Sample
- 20254 observations
- Configuration
- Claude Opus 5
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.10
Confirmed task success · rank 2 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- 0.07 to 0.13
- Sample
- 14987 observations
- Configuration
- Claude Opus 5
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.18
Praise over complaint · rank 3 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- 0.13 to 0.23
- Sample
- 8011 observations
- Configuration
- Claude Opus 5
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.17
Praise over complaint · rank 3 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- 0.11 to 0.22
- Sample
- 5796 observations
- Configuration
- Claude Opus 5
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.11
Steerability · rank 1 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- 0.08 to 0.13
- Sample
- 25690 observations
- Configuration
- Claude Opus 5
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.08
Steerability · rank 1 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- 0.05 to 0.11
- Sample
- 17700 observations
- Configuration
- Claude Opus 5
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- 0.00 to 0.00
- Sample
- 3160049 observations
- Configuration
- Claude Opus 5
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- 0.00 to 0.00
- Sample
- 2704226 observations
- Configuration
- Claude Opus 5
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,490
Overall · rank 1 of 44
- Unit
- Arena rating, higher is better
- Range
- 1,483 to 1,497
- Sample
- 8794 votes
- Configuration
- Claude Opus 5 (high)
- Measured
- 13 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,477
Overall · rank 7 of 177
- Unit
- Arena rating, higher is better
- Range
- 1,474 to 1,480
- Sample
- 54398 votes
- Configuration
- Claude Opus 5 (high)
- Measured
- 25 Sep 2026
- Not shown
- Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,488
Overall · rank 2 of 177
- Unit
- Arena rating, higher is better
- Range
- 1,484 to 1,493
- Sample
- 26471 votes
- Configuration
- Claude Opus 5 (max)
- Measured
- 25 Sep 2026
- Not shown
- Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,486
Business, management and finance · rank 6 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,479 to 1,493
- Sample
- 10655 votes
- Configuration
- Claude Opus 5 (high)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,482
Business, management and finance · rank 7 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,473 to 1,491
- Sample
- 5132 votes
- Configuration
- Claude Opus 5 (max)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,472
Creative writing · rank 6 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,465 to 1,479
- Sample
- 11902 votes
- Configuration
- Claude Opus 5 (high)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,470
Creative writing · rank 6 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,461 to 1,479
- Sample
- 5870 votes
- Configuration
- Claude Opus 5 (max)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,543
Expert prompts · rank 1 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,534 to 1,551
- Sample
- 6288 votes
- Configuration
- Claude Opus 5 (high)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,527
Expert prompts · rank 2 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,516 to 1,539
- Sample
- 2915 votes
- Configuration
- Claude Opus 5 (max)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,497
Instruction following · rank 2 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,491 to 1,503
- Sample
- 19769 votes
- Configuration
- Claude Opus 5 (high)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,493
Instruction following · rank 3 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,485 to 1,500
- Sample
- 9392 votes
- Configuration
- Claude Opus 5 (max)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,491
Overall · rank 5 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,487 to 1,495
- Sample
- 55063 votes
- Configuration
- Claude Opus 5 (high)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,488
Overall · rank 7 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,483 to 1,493
- Sample
- 26760 votes
- Configuration
- Claude Opus 5 (max)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,482
Writing, literature and language · rank 5 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,476 to 1,489
- Sample
- 15394 votes
- Configuration
- Claude Opus 5 (high)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,480
Writing, literature and language · rank 5 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,473 to 1,488
- Sample
- 7534 votes
- Configuration
- Claude Opus 5 (max)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
32.0%
Consistency (pass^4) · rank 2 of 21
- Unit
- % of tasks, higher is better
- Configuration
- Claude Opus 5 (max reasoning, tau2)
- Measured
- 3 Aug 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
48.7%
Task success (pass^1) · rank 2 of 21
- Unit
- % of tasks, higher is better
- Configuration
- Claude Opus 5 (max reasoning, tau2)
- Measured
- 3 Aug 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.