Models / qwen3-max-preview
Alibabaqwen3-max-preview
14 published results from 2 sources. Each card shows where the number comes from and what it does not measure. The overall leaderboard combines them; here each stands alone.
- Provider
- Alibaba
- Sources
- 2
- Our benchmarks
- 0
- Price per million tokens
- Not listed on OpenRouter
Reported by others
1,442
Overall · rank 58 of 177
- Unit
- Arena rating, higher is better
- Range
- 1,435 to 1,448
- Sample
- 7517 votes
- Configuration
- qwen3-max-preview
- Measured
- 25 Sep 2026
- Not shown
- Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,445
Business, management and finance · rank 52 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,436 to 1,454
- Sample
- 5126 votes
- Configuration
- qwen3-max-preview
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,393
Creative writing · rank 87 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,383 to 1,403
- Sample
- 3689 votes
- Configuration
- qwen3-max-preview
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,464
Expert prompts · rank 49 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,448 to 1,480
- Sample
- 1322 votes
- Configuration
- qwen3-max-preview
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,424
Instruction following · rank 77 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,417 to 1,432
- Sample
- 7258 votes
- Configuration
- qwen3-max-preview
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,435
Overall · rank 85 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,430 to 1,439
- Sample
- 27235 votes
- Configuration
- qwen3-max-preview
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,405
Writing, literature and language · rank 89 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,398 to 1,413
- Sample
- 6133 votes
- Configuration
- qwen3-max-preview
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
12.6
Artificial Analysis Intelligence Index · rank 243 of 314
- Unit
- index score, higher is better
- Configuration
- qwen3-max-preview
- Measured
- 1 Oct 2026
- Not shown
- Not business work, and a blend: read the parts for any one task.
76.4%
GPQA Diamond · rank 181 of 276
- Unit
- % of questions, higher is better
- Configuration
- qwen3-max-preview
- Measured
- 1 Oct 2026
- Not shown
- Not applied work; multiple-choice science questions.
10.1%
Humanity's Last Exam · rank 228 of 314
- Unit
- % of questions, higher is better
- Configuration
- qwen3-max-preview
- Measured
- 1 Oct 2026
- Not shown
- Not everyday work; academic questions at the edge of expertise.
48.0%
IFBench · rank 124 of 215
- Unit
- % of instructions, higher is better
- Configuration
- qwen3-max-preview
- Measured
- 1 Oct 2026
- Not shown
- Not judgement about what an instruction meant.
43.0%
Long-context reasoning (AA-LCR) · rank 250 of 309
- Unit
- % of questions, higher is better
- Configuration
- qwen3-max-preview
- Measured
- 1 Oct 2026
- Not shown
- Not retrieval over your own document store.
19.7%
Terminal-Bench Hard · rank 131 of 213
- Unit
- % of tasks, higher is better
- Configuration
- qwen3-max-preview
- Measured
- 1 Oct 2026
- Not shown
- Not other harnesses; Artificial Analysis no longer runs it on new models.
32.7%
Τ²-bench telecom · rank 168 of 214
- Unit
- % of tasks, higher is better
- Configuration
- qwen3-max-preview
- Measured
- 1 Oct 2026
- Not shown
- Not your policies or systems; no longer run on new models.
Compare qwen3-max-preview with