Models / Qwen3-Max-Thinking

Alibaba

Qwen3-Max-Thinking

Early evidence, 7 results from 1 source, on 1 use case. Qwen3-Max-Thinking is strongest at reasoning and knowledge; there is too little evidence yet to recommend it either way. The cheapest model of similar intelligence is Solar Pro 4, at $0.16 per million tokens against $1.56.

28th most intelligent of the 38 models priced $1–3 per million tokens. Among models available today, mid-field on reasoning and knowledge, the one use case measured. Results for agents and tool use are pending: those sources last updated before this model was released.

Results as of 8 October 2026, from 7 results on 1 source; prices checked 8 Oct 2026.

Price per million tokens
$0.78 in · $3.90 out
Typical host on OpenRouter, 8 Oct 2026 · compare prices
Speed
No speed yet
Artificial Analysis lists it but has not published a speed for it yet
Intelligence Index
21.3
Artificial Analysis
Spring Prompt overall
Not ranked yet
needs our benchmarks and two groups of results
Developer
Alibaba, based in China
Weights not published
Released
26 Jan 2026 (Artificial Analysis)
262,144 tokens of context

Where it stands among the models you could choose

Ranked only among models available today (those you can call through OpenRouter), in the groups people choose between: the same price band, similar intelligence, the same speed, developers based in the same place. Retired models are left out.

AmongIntelligence IndexSpring Prompt overallPrice, blendedSpeed
Models available today 90th of 198Claude Opus 5.5 – $1.56 blended, 64th percentilecheapest: Mistral Nemo –
Models priced $1–3 per million tokens 28th of 38Muse Spark 1.3 – sets this group –
Models with similar intelligence sets this group – $1.56 blended, 67th percentilecheapest: DeepSeek-V4-Flash (0423) –
Models from developers based in China 40th of 86MiMo-V2.6-Pro – $1.56 blended, 84th percentilecheapest: Qwen3.5-9B –

The name under each rank is the leader of that group; green is the top quarter of the group and red the bottom quarter, by rank. Price is shown as a percentile: the share of the group that costs less, so lower is cheaper. "Sets this group" marks the measure the group is defined by. Spring Prompt overall is our score out of 100 across every source (how it works). Price is the developer's list price (or the typical OpenRouter host where we haven't read one) for three input tokens to one output; intelligence and speed are from Artificial Analysis (data sourced from Artificial Analysis); "similar intelligence" means within 4 points on its Intelligence Index. Developer locations are where each company is based, not where a model is served. Groups under 3 models are not ranked.

How good is it, and for what?

Each use case ranks the available models its benchmarks measured, then those priced $1–3 per million tokens. The bar shows its percentile in that field, best to the right. Open a row for the results behind it.

Product listingsTurning a sparse product feed and photos into listings that can go live Not measured
  • CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is Not measured
  • DeckBench: has not measured this model.
User surveysPlanning a user survey and reading its results without being misled Not measured
  • SurveyBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending Not measured
  • ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls PendingReleased after these sources last updated
  • tau2-bench: has not measured this model.
  • Berkeley Function Calling Leaderboard (BFCL) V4: pending. Its latest results are dated 16 Dec 2025, before this model was released.
  • OpenHands Index: has not measured this model.
  • Microsoft STATE-Bench: has not measured this model.
  • Artificial Analysis: has not measured this model.
  • Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects Not measured
  • APEX-Agents: has not measured this model.
  • GDP.pdf: has not measured this model.
  • Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge 81st of 198 available · 25th of 38 at its price
BenchmarkQwen3-Max-ThinkingRank among availableBest available
Artificial AnalysisArtificial Analysis: Artificial Analysis Intelligence Index · reported 21.3 90th of 198 Claude Opus 5.5 57.6
Artificial AnalysisArtificial Analysis: Humanity's Last Exam · reported 28.0% 76th of 197 Claude Opus 5.5 61.4%
Artificial AnalysisArtificial Analysis: GPQA Diamond · reported 86.1% 64th of 187 GPT-6 Astra 96.3%
WritingWhat people prefer in blind comparisons, and judged writing quality Not measured
  • Arena (formerly LMArena): has not measured this model.
  • UGI Leaderboard: has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers Not measured
  • Vectara Hallucination Leaderboard: has not measured this model.
  • Arena (formerly LMArena): has not measured this model.
  • SimpleQA Verified (Epoch AI): has not measured this model.
Speed under pressureGood decisions against a real clock (fast chess) Not measured
  • BulletBench: has not measured this model.

Nearest alternatives

  • Cheaper at similar intelligenceSolar Pro 4$0.16 blended · Intelligence Index 28.2
  • Alibaba's newer, at least as capable modelQwen3.8-Max (0902)$3.00 blended · Intelligence Index 45.4
  • Open weights at similar intelligenceMiMo-V2.6-Flash$0.18 blended · Intelligence Index 37.9

Against Qwen3-Max-Thinking at $1.56 blended, Intelligence Index 21.3. "Similar intelligence" means within 4 points or better.

Where to run it

1 host on OpenRouter at the standard service tier, with a median of $1.56 blended per million tokens. Cheapest first.

HostQuantisationInputOutputBlended
Alibabanot stated$0.78$3.90$1.56

Prices per million tokens from OpenRouter's public endpoint list, checked 8 Oct 2026.

Against the alternatives

The models you would most likely weigh it against: the leaders of the groups above.

Scroll sideways to see every alternative.

Qwen3-Max-ThinkingMuse Spark 1.3
most intelligent at $1–3 per million tokens
DeepSeek-V4-Flash (0423)
cheapest with similar intelligence
Claude Opus 5.5
most intelligent available today
MiMo-V2.6-Pro
most intelligent from China
Price per million tokens $1.56 $2.00$0.18$8.00$0.54
Tokens a second – 114–9739
Intelligence Index 21.3 48.124.457.646.3
Spring Prompt overall – 65–85–
Rank among available models, by use case
Product listings – 9th–6th–
Decks from an analysis – 14th–5th–
User surveys – ––2nd–
Marketing planning – 12th–5th–
Agents and tool use pending 9th52nd2nd15th
Professional work – 8th–3rd–
Reasoning and knowledge 81st of 198 8th64th1st10th
Writing – 5th62nd4th20th
Sticking to the facts – 11th83rd2nd8th
Speed under pressure – 19th–––

Shaded figures are better than Qwen3-Max-Thinking's. A dash means no result.

Where it has been measured

Every result

7 published results from 1 source, each in the source's own units, with the configuration that produced it.

Artificial Analysis · reported by Artificial Analysis · 7 results
MeasureValueRankConfigurationDated
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 21.3 191 of 407 Qwen3-Max-Thinking 8 Oct 2026
Artificial Analysis: GPQA Diamond% of questions, higher is better 86.1% 111 of 359 Qwen3-Max-Thinking 8 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 28.0% 150 of 405 Qwen3-Max-Thinking 8 Oct 2026
Artificial Analysis: IFBench% of instructions, higher is better 70.7% 59 of 288 Qwen3-Max-Thinking 8 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 74.3% 146 of 393 Qwen3-Max-Thinking 8 Oct 2026
Artificial Analysis: Terminal-Bench Hard% of tasks, higher is better 24.2% 132 of 282 Qwen3-Max-Thinking 8 Oct 2026
Artificial Analysis: τ²-bench telecom% of tasks, higher is better 83.6% 91 of 286 Qwen3-Max-Thinking 8 Oct 2026

Not shown: Not business work, and a blend: read the parts for any one task.