Models / GPT-4o-mini (2024-07-18)
OpenAIGPT-4o-mini (2024-07-18)
6 published results from 1 source. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.
- Provider
- OpenAI
- Sources
- 1
- Our benchmarks
- 0
- Price
- Not yet published
Reported by others
1,311
Business, management and finance · rank 225 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,303 to 1,319
- Sample
- 7765 votes
- Configuration
- GPT-4o-mini (2024-07-18)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,294
Creative writing · rank 211 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,287 to 1,301
- Sample
- 10484 votes
- Configuration
- GPT-4o-mini (2024-07-18)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,285
Expert prompts · rank 242 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,274 to 1,296
- Sample
- 3549 votes
- Configuration
- GPT-4o-mini (2024-07-18)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,293
Instruction following · rank 248 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,288 to 1,298
- Sample
- 26705 votes
- Configuration
- GPT-4o-mini (2024-07-18)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,318
Overall · rank 244 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,314 to 1,321
- Sample
- 68697 votes
- Configuration
- GPT-4o-mini (2024-07-18)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,314
Writing, literature and language · rank 212 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,308 to 1,319
- Sample
- 18531 votes
- Configuration
- GPT-4o-mini (2024-07-18)
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.