1,397
Business, management and finance · rank 108 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,375 to 1,418
- Sample
- 718 votes
- Configuration
- DeepSeek V3.1 Terminus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,423
Business, management and finance · rank 64 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,401 to 1,445
- Sample
- 687 votes
- Configuration
- DeepSeek V3.1 Terminus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,405
Creative writing · rank 52 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,378 to 1,432
- Sample
- 446 votes
- Configuration
- DeepSeek V3.1 Terminus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,388
Creative writing · rank 73 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,360 to 1,417
- Sample
- 431 votes
- Configuration
- DeepSeek V3.1 Terminus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,394
Instruction following · rank 107 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,377 to 1,412
- Sample
- 1016 votes
- Configuration
- DeepSeek V3.1 Terminus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,420
Instruction following · rank 63 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,400 to 1,440
- Sample
- 849 votes
- Configuration
- DeepSeek V3.1 Terminus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,415
Overall · rank 110 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,405 to 1,424
- Sample
- 3600 votes
- Configuration
- DeepSeek V3.1 Terminus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,417
Overall · rank 106 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,407 to 1,427
- Sample
- 3351 votes
- Configuration
- DeepSeek V3.1 Terminus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,408
Writing, literature and language · rank 70 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,387 to 1,428
- Sample
- 787 votes
- Configuration
- DeepSeek V3.1 Terminus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,411
Writing, literature and language · rank 67 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,390 to 1,433
- Sample
- 752 votes
- Configuration
- DeepSeek V3.1 Terminus
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
34.0%
Requested-length error · rank 269 of 370
- Unit
- % off the requested word count, lower is better
- Configuration
- DeepSeek V3.1 Terminus (no reasoning)
- Measured
- 2 Oct 2025
- Not shown
- Not other format limits such as character counts or bullet counts.
32.0%
Requested-length error · rank 263 of 370
- Unit
- % off the requested word count, lower is better
- Configuration
- DeepSeek V3.1 Terminus
- Measured
- 2 Oct 2025
- Not shown
- Not other format limits such as character counts or bullet counts.
0.36
Style adherence · rank 127 of 370
- Unit
- score from 0 to 1, higher is better
- Configuration
- DeepSeek V3.1 Terminus (no reasoning)
- Measured
- 2 Oct 2025
- Not shown
- Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
0.37
Style adherence · rank 88 of 370
- Unit
- score from 0 to 1, higher is better
- Configuration
- DeepSeek V3.1 Terminus
- Measured
- 2 Oct 2025
- Not shown
- Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
52.0
Writing score · rank 146 of 370
- Unit
- score out of 100, higher is better
- Configuration
- DeepSeek V3.1 Terminus (no reasoning)
- Measured
- 2 Oct 2025
- Not shown
- Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
54.0
Writing score · rank 133 of 370
- Unit
- score out of 100, higher is better
- Configuration
- DeepSeek V3.1 Terminus
- Measured
- 2 Oct 2025
- Not shown
- Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.