Models / gpt-4-1106-preview

OpenAI

gpt-4-1106-preview

6 published results from 1 source. Each card shows where the number comes from and what it does not measure. The overall leaderboard combines them; here each stands alone.

Provider
OpenAI
Sources
1
Our benchmarks
0
Price per million tokens
Not listed on OpenRouter

Reported by others

1,279
Business, management and finance · rank 288 of 402
Unit
Arena rating, higher is better
Range
1,271 to 1,288
Sample
9573 votes
Configuration
gpt-4-1106-preview
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,286
Creative writing · rank 249 of 407
Unit
Arena rating, higher is better
Range
1,278 to 1,294
Sample
15576 votes
Configuration
gpt-4-1106-preview
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,290
Expert prompts · rank 268 of 359
Unit
Arena rating, higher is better
Range
1,278 to 1,302
Sample
4240 votes
Configuration
gpt-4-1106-preview
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,297
Instruction following · rank 261 of 409
Unit
Arena rating, higher is better
Range
1,291 to 1,302
Sample
34416 votes
Configuration
gpt-4-1106-preview
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,313
Overall · rank 266 of 409
Unit
Arena rating, higher is better
Range
1,309 to 1,317
Sample
100105 votes
Configuration
gpt-4-1106-preview
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,308
Writing, literature and language · rank 240 of 408
Unit
Arena rating, higher is better
Range
1,301 to 1,314
Sample
24682 votes
Configuration
gpt-4-1106-preview
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.