Models / qwen1.5-4b-chat

Alibaba

qwen1.5-4b-chat

6 published results from 1 source. Each card shows where the number comes from and what it does not measure. The overall leaderboard combines them; here each stands alone.

Provider
Alibaba
Sources
1
Our benchmarks
0
Price per million tokens
Not listed on OpenRouter

Reported by others

1,084
Business, management and finance · rank 367 of 402
Unit
Arena rating, higher is better
Range
1,060 to 1,108
Sample
738 votes
Configuration
qwen1.5-4b-chat
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,050
Creative writing · rank 387 of 407
Unit
Arena rating, higher is better
Range
1,031 to 1,069
Sample
1080 votes
Configuration
qwen1.5-4b-chat
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,108
Expert prompts · rank 339 of 359
Unit
Arena rating, higher is better
Range
1,077 to 1,139
Sample
356 votes
Configuration
qwen1.5-4b-chat
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,070
Instruction following · rank 384 of 409
Unit
Arena rating, higher is better
Range
1,057 to 1,083
Sample
2636 votes
Configuration
qwen1.5-4b-chat
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,091
Overall · rank 393 of 409
Unit
Arena rating, higher is better
Range
1,082 to 1,100
Sample
7597 votes
Configuration
qwen1.5-4b-chat
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,077
Writing, literature and language · rank 380 of 408
Unit
Arena rating, higher is better
Range
1,062 to 1,092
Sample
1895 votes
Configuration
qwen1.5-4b-chat
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.