Models / Qwen3.8 Max (0902)

Alibaba

Qwen3.8 Max (0902)

35 published results from 5 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Alibaba
Sources
5
Our benchmarks
3
Price
Not yet published

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
0
Ladder Elo · rank 12 of 16
Unit
ladder Elo, higher is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
7.4 s
Median move time · rank 11 of 16
Unit
milliseconds, lower is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · CatalogBench
92.5%
Content quality · rank 17 of 18
Unit
% of checks, higher is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
3.6%
Failed outputs · rank 17 of 18
Unit
% of products, lower is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
0.0%
Publish-ready listings · rank 15 of 18
Unit
% of products, higher is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Reliably publish-ready · rank 10 of 18
Unit
% of products, higher is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
9.09
Unsupported claims · rank 16 of 18
Unit
claims per product, lower is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
99.4%
Channel compliance · rank 12 of 18
Unit
% of products, higher is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Conflicts caught · rank 1 of 18
Unit
% of conflicts, higher is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
93.5%
Content quality · rank 17 of 18
Unit
% of checks, higher is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
$0.0256
Cost per product · rank 13 of 18
Unit
US dollars, lower is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · CatalogBench
91.0%
Decision accuracy · rank 15 of 18
Unit
% of decisions, higher is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
3.6%
Failed outputs · rank 17 of 18
Unit
% of products, lower is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
90.3%
Field accuracy · rank 17 of 18
Unit
% of missing fields, higher is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
2.7%
Invented values · rank 14 of 18
Unit
% of filled values, lower is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
28.0%
Publish-ready listings · rank 14 of 18
Unit
% of products, higher is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
10.7%
Reliably publish-ready · rank 14 of 18
Unit
% of products, higher is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
1.62
Unsupported claims · rank 11 of 18
Unit
claims per product, lower is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
30 Sep 2026
Failures
reply cut off at the token limit: 6 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · ROASBench
$0.73
Cost of a run · rank 13 of 17
Unit
US dollars, lower is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
29 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
2
Months over budget · rank 15 of 17
Unit
months of 12, lower is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
25.4
Overall score · rank 13 of 17
Unit
score out of 100, higher is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
£0.09
Return on ad spend · rank 13 of 17
Unit
profit per £1 spent, higher is better
Configuration
Qwen3.8 Max (0902), provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

Reported by others

0.05
Confirmed task success · rank 10 of 46
Unit
IPS effect estimate, higher is better
Range
0.04 to 0.07
Sample
32832 observations
Configuration
Qwen3.8 Max (0902)
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.04
Praise over complaint · rank 10 of 46
Unit
IPS effect estimate, higher is better
Range
0.01 to 0.06
Sample
12602 observations
Configuration
Qwen3.8 Max (0902)
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.02
Steerability · rank 11 of 46
Unit
IPS effect estimate, higher is better
Range
0.01 to 0.03
Sample
39809 observations
Configuration
Qwen3.8 Max (0902)
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.01
Tool grounding · rank 39 of 46
Unit
IPS effect estimate, higher is better
Range
-0.01 to -0.01
Sample
3755021 observations
Configuration
Qwen3.8 Max (0902)
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,471
Overall · rank 10 of 177
Unit
Arena rating, higher is better
Range
1,466 to 1,475
Sample
21287 votes
Configuration
Qwen3.8 Max (0902)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,480
Business, management and finance · rank 7 of 402
Unit
Arena rating, higher is better
Range
1,470 to 1,490
Sample
4132 votes
Configuration
Qwen3.8 Max (0902)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,467
Creative writing · rank 6 of 407
Unit
Arena rating, higher is better
Range
1,457 to 1,478
Sample
4156 votes
Configuration
Qwen3.8 Max (0902)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,510
Expert prompts · rank 8 of 359
Unit
Arena rating, higher is better
Range
1,498 to 1,522
Sample
2436 votes
Configuration
Qwen3.8 Max (0902)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,472
Instruction following · rank 13 of 409
Unit
Arena rating, higher is better
Range
1,465 to 1,480
Sample
7762 votes
Configuration
Qwen3.8 Max (0902)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,479
Overall · rank 13 of 409
Unit
Arena rating, higher is better
Range
1,474 to 1,485
Sample
21561 votes
Configuration
Qwen3.8 Max (0902)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,472
Writing, literature and language · rank 8 of 408
Unit
Arena rating, higher is better
Range
1,463 to 1,480
Sample
5502 votes
Configuration
Qwen3.8 Max (0902)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by tau2-bench
35.1%
Consistency (pass^4) · rank 1 of 21
Unit
% of tasks, higher is better
Configuration
qwen3.8-max-0902
Measured
3 Aug 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
55.2%
Task success (pass^1) · rank 1 of 21
Unit
% of tasks, higher is better
Configuration
qwen3.8-max-0902
Measured
3 Aug 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.

Compare Qwen3.8 Max (0902) with