Models / Qwen3.8-Max

Alibaba

Qwen3.8-Max

Early evidence, 10 results from 2 sources, on 3 use cases. Qwen3.8-Max is best for reasoning and knowledge, and agents and tool use; it is not on sale through OpenRouter, so it is here for reference.

Superseded. Qwen3.8-Max (0902) is Alibaba's newer model in this line, at $3.00 blended: see Qwen3.8-Max (0902) or compare the two.

Compared with the models available today (it is not on sale itself), in the top quarter (by percentile) for reasoning and knowledge, and agents and tool use.

Results as of 8 October 2026, from 10 results on 2 sources; prices checked 8 Oct 2026.

Price per million tokens
Not available through OpenRouter
Speed
No speed yet
Artificial Analysis lists it but has not published a speed for it yet
Intelligence Index
40.2
Artificial Analysis, at 0803
Spring Prompt overall
Not ranked yet
needs our benchmarks and two groups of results
Developer
Alibaba, based in China
Weights not published
Released
3 Aug 2026 (Artificial Analysis)

Where it stands among the models you could choose

Ranked only among models available today (those you can call through OpenRouter), in the groups people choose between: the same price band, similar intelligence, the same speed, developers based in the same place. Retired models are left out.

AmongIntelligence IndexSpring Prompt overallPrice, blendedSpeed
Models available today 23rd of 199Claude Opus 5.5 – – –
Models from developers based in China 6th of 87MiMo-V2.6-Pro – – –

The name under each rank is the leader of that group; green is the top quarter of the group and red the bottom quarter, by rank. Price is shown as a percentile: the share of the group that costs less, so lower is cheaper. "Sets this group" marks the measure the group is defined by. Spring Prompt overall is our score out of 100 across every source (how it works). Price is the developer's list price (or the typical OpenRouter host where we haven't read one) for three input tokens to one output; intelligence and speed are from Artificial Analysis (data sourced from Artificial Analysis); "similar intelligence" means within 4 points on its Intelligence Index. Developer locations are where each company is based, not where a model is served. Groups under 3 models are not ranked.

How good is it, and for what?

Each use case ranks the available models its benchmarks measured. The bar shows its percentile in that field, best to the right. Open a row for the results behind it.

Product listingsTurning a sparse product feed and photos into listings that can go live Not measured
  • CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is Not measured
  • DeckBench: has not measured this model.
User surveysPlanning a user survey and reading its results without being misled Not measured
  • SurveyBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending Not measured
  • ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls 18th of 165 available
BenchmarkQwen3.8-MaxRank among availableBest available
Artificial AnalysisArtificial Analysis: Terminal-Bench 4.0 · reported 18.7% 24th of 106 Claude Sonnet 5.5 63.6%
Artificial AnalysisArtificial Analysis: Terminal-Bench 2.1 · reported 81.3% 20th of 116 Claude Fable 5.1 91.4%
Artificial AnalysisArtificial Analysis: τ-bench banking · reported 51.3% 1st of 114 This model
  • tau2-bench: pending. Its latest results are dated 26 May 2026, before this model was released.
  • Berkeley Function Calling Leaderboard (BFCL) V4: pending. Its latest results are dated 16 Dec 2025, before this model was released.
  • OpenHands Index: pending. Its latest results are dated 30 Jun 2026, before this model was released.
  • Microsoft STATE-Bench: pending. Its latest results are dated 29 May 2026, before this model was released.
  • Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects Not measured
  • APEX-Agents: has not measured this model.
  • GDP.pdf: has not measured this model.
  • Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge 23rd of 199 available
BenchmarkQwen3.8-MaxRank among availableBest available
Artificial AnalysisArtificial Analysis: Artificial Analysis Intelligence Index · reported 40.2 23rd of 199 Claude Opus 5.5 57.6
Artificial AnalysisArtificial Analysis: Humanity's Last Exam · reported 43.0% 25th of 198 Claude Opus 5.5 61.4%
Artificial AnalysisArtificial Analysis: GPQA Diamond · reported 92.7% 18th of 188 GPT-6 Astra 96.3%
WritingWhat people prefer in blind comparisons, and judged writing quality Not measured
  • Arena (formerly LMArena): has not measured this model.
  • UGI Leaderboard: has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers 66th of 153 available
BenchmarkQwen3.8-MaxRank among availableBest available
SimpleQA Verified (Epoch AI)SimpleQA Verified: correct answers · reported 45.8% 35th of 74 GPT-6 Astra 75.6%
  • Vectara Hallucination Leaderboard: has not measured this model.
  • Arena (formerly LMArena): has not measured this model.
Speed under pressureGood decisions against a real clock (fast chess) Not measured
  • BulletBench: has not measured this model.

Nearest alternatives

Against Qwen3.8-Max, Intelligence Index 40.2. "Similar intelligence" means within 4 points or better.

Against the alternatives

The models you would most likely weigh it against: the leaders of the groups above.

Scroll sideways to see every alternative.

Qwen3.8-MaxClaude Opus 5.5
most intelligent available today
MiMo-V2.6-Pro
most intelligent from China
Qwen3.8-Max (0902)
Alibaba's best other model
GPT-6.1 Sol
near the top overall
Price per million tokens – $8.00$0.54$3.00$4.00
Tokens a second – 97393755
Intelligence Index 40.2 57.646.345.451.8
Spring Prompt overall – 85–4881
Rank among available models, by use case
Product listings – 6th–17th2nd
Decks from an analysis – 5th–11th3rd
User surveys – 2nd––1st
Marketing planning – 5th–14th–
Agents and tool use 18th of 165 2nd14th7th3rd
Professional work – 3rd–18th9th
Reasoning and knowledge 23rd of 199 1st10th17th5th
Writing – 4th20th16th15th
Sticking to the facts 66th of 153 2nd8th37th18th
Speed under pressure – ––––

Shaded figures are better than Qwen3.8-Max's. A dash means no result.

Where it has been measured

Every result

10 published results from 2 sources, each in the source's own units, with the configuration that produced it.

Artificial Analysis · reported by Artificial Analysis · 9 results
MeasureValueRankConfigurationDated
Artificial Analysis: Artificial Analysis Coding Indexindex score, higher is better 71.8 41 of 188 Qwen3.8-Max (0803) 8 Oct 2026
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 40.2 56 of 407 Qwen3.8-Max (0803) 8 Oct 2026
Artificial Analysis: GPQA Diamond% of questions, higher is better 92.7% 32 of 359 Qwen3.8-Max (0803) 8 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 43.0% 54 of 405 Qwen3.8-Max (0803) 8 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 78.3% 100 of 393 Qwen3.8-Max (0803) 8 Oct 2026
Artificial Analysis: SciCode% of problems, higher is better 53.2% 71 of 191 Qwen3.8-Max (0803) 8 Oct 2026
Artificial Analysis: Terminal-Bench 2.1% of tasks, higher is better 81.3% 41 of 187 Qwen3.8-Max (0803) 8 Oct 2026
Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better 18.7% 59 of 187 Qwen3.8-Max (0803) 8 Oct 2026
Artificial Analysis: τ-bench banking% of tasks, higher is better 51.3% 1 of 175 Qwen3.8-Max (0803) 8 Oct 2026

Not shown: Not your codebase or tools.

SimpleQA Verified (Epoch AI) · reported by SimpleQA Verified (Epoch AI) · 1 result
MeasureValueRankConfigurationDated
SimpleQA Verified: correct answers% of questions, higher is better 45.8% 36 of 85 Qwen3.8-Max (0803, xhigh reasoning) 27 Aug 2026

Not shown: Not answers grounded in your documents; tests what the model remembers.