Models / Qwen3.8 2.4T A95B

Alibaba

Qwen3.8 2.4T A95B

11th most intelligent of the 34 models priced $3–10 per million tokens. Among models available today, in the top quarter for reasoning and knowledge, and agents and tool use.

Price per million tokens
$2.00 in · $6.00 out
Alibaba's list price, 7 Oct 2026 · compare prices
Speed
38 tokens a second
Artificial Analysis, on its usual API
Intelligence Index
39.9
Artificial Analysis
Spring Prompt overall
Not ranked yet
needs our benchmarks and two groups of results
Developer
Alibaba, based in China
Weights published
Released
Not recorded
1,048,576 tokens of context

Where it stands among the models you could choose

Ranked only among models available today (those you can call through OpenRouter), in the groups people choose between: the same price band, similar intelligence, the same speed, developers based in the same place, open weights. Retired models are left out.

AmongIntelligence IndexSpring Prompt overallPriceSpeed
Models available today 22nd of 195Claude Opus 5.5 – 170th of 221Mistral Nemo 63rd of 63Trinity Large Thinking
Models priced $3–10 per million tokens 11th of 34Claude Opus 5.5 – the group 13th of 13Mistral Medium 3.5
Models with similar intelligence the group – 11th of 19MiMo-V2.6-Flash 10th of 10DeepSeek V4.1 Flash
Models writing under 50 tokens a second 5th of 9MiMo-V2.6-Pro – 7th of 9Phi 4 the group
Models from developers based in China 6th of 84MiMo-V2.6-Pro – 94th of 96Qwen3.5-9B 22nd of 22DeepSeek V4 Flash Vision Exp
Open-weights modelsincluding 1 with weights announced but not yet published 5th of 108MiMo-V2.6-Pro – 114th of 117Mistral Nemo 41st of 41Trinity Large Thinking

The name under each rank is the leader of that group. Price is the developer's list price (or the typical OpenRouter host where we haven't read one) for three input tokens to one output; intelligence and speed are from Artificial Analysis (data sourced from Artificial Analysis); "similar intelligence" means within 4 points on its Intelligence Index. Developer locations are where each company is based, not where a model is served. Groups under 3 models are not ranked.

How good is it, and for what?

Each use case ranks the available models its benchmarks measured, then those priced $3–10 per million tokens. The bar shows where it falls in that field, best to the right. Open a row for the results behind it.

Product listingsTurning a sparse product feed and photos into listings that can go live Not measured
  • CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is Not measured
  • DeckBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending Not measured
  • ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls 23rd of 161 available · 10th of 30 at its price
BenchmarkQwen3.8 2.4T A95BRank among availableBest available
Artificial AnalysisArtificial Analysis: Terminal-Bench 4.0 · reported 11.1% 32nd of 102 Claude Sonnet 5.5 63.6%
Artificial AnalysisArtificial Analysis: Terminal-Bench 2.1 · reported 82.0% 18th of 113 Claude Fable 5.1 91.4%
Artificial AnalysisArtificial Analysis: τ-bench banking · reported 49.1% 4th of 110 Grok 4.6 50.7%
  • tau2-bench: has not measured this model.
  • Berkeley Function Calling Leaderboard (BFCL) V4: has not measured this model.
  • OpenHands Index: has not measured this model.
  • Microsoft STATE-Bench: has not measured this model.
  • Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects Not measured
  • APEX-Agents: has not measured this model.
  • GDP.pdf: has not measured this model.
  • Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge 21st of 195 available · 11th of 34 at its price
BenchmarkQwen3.8 2.4T A95BRank among availableBest available
Artificial AnalysisArtificial Analysis: Artificial Analysis Intelligence Index · reported 39.9 22nd of 195 Claude Opus 5.5 57.6
Artificial AnalysisArtificial Analysis: Humanity's Last Exam · reported 42.4% 28th of 194 Claude Opus 5.5 61.4%
Artificial AnalysisArtificial Analysis: GPQA Diamond · reported 93.5% 10th of 185 GPT-6 Astra 96.3%
WritingWhat people prefer in blind comparisons, and judged writing quality Not measured
  • Arena (formerly LMArena): has not measured this model.
  • UGI Leaderboard: has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers Not measured
  • Vectara Hallucination Leaderboard: has not measured this model.
  • Arena (formerly LMArena): has not measured this model.
  • SimpleQA Verified (Epoch AI): has not measured this model.
Speed under pressureGood decisions against a real clock (fast chess) Not measured
  • BulletBench: has not measured this model.

Against the alternatives

The models you would most likely weigh it against: the leaders of the groups above.

Qwen3.8 2.4T A95BClaude Opus 5.5
most intelligent at $3–10 per million tokens
MiMo-V2.6-Flash
cheapest with similar intelligence
DeepSeek V4.1 Flash
fastest with similar intelligence
MiMo-V2.6-Pro
most intelligent from China
Price per million tokens $3.00 $8.00$0.18$0.52$0.54
Tokens a second 38 975221640
Intelligence Index 39.9 57.637.939.546.3
Spring Prompt overall – 82–58–
Rank among available models, by use case
Product listings – 6th–7th–
Decks from an analysis – 5th–––
Marketing planning – 5th–––
Agents and tool use 23rd of 161 2nd29th20th15th
Professional work – 3rd–––
Reasoning and knowledge 21st of 195 1st44th34th10th
Writing – 4th50th27th18th
Sticking to the facts – 6th40th19th9th
Speed under pressure – ––––

Shaded figures are better than Qwen3.8 2.4T A95B's. A dash means no result.

Where it has been measured

Every result

11 published results from 1 source, each in the source's own units, with the configuration that produced it.

Artificial Analysis · reported by Artificial Analysis · 11 results
MeasureValueRankConfigurationDated
Artificial Analysis: Artificial Analysis Coding Indexindex score, higher is better 71.9 40 of 188 Qwen3.8 2.4T A95B 7 Oct 2026
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 39.9 55 of 402 Qwen3.8 2.4T A95B 7 Oct 2026
Artificial Analysis: GPQA Diamond% of questions, higher is better 93.5% 14 of 359 Qwen3.8 2.4T A95B 7 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 42.4% 59 of 400 Qwen3.8 2.4T A95B 7 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 80.3% 61 of 388 Qwen3.8 2.4T A95B 7 Oct 2026
Artificial Analysis: Output speedtokens per second, higher is better 38.0 123 of 128 Qwen3.8 2.4T A95B 7 Oct 2026
Artificial Analysis: SciCode% of problems, higher is better 54.1% 62 of 186 Qwen3.8 2.4T A95B 7 Oct 2026
Artificial Analysis: Terminal-Bench 2.1% of tasks, higher is better 82.0% 39 of 187 Qwen3.8 2.4T A95B 7 Oct 2026
Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better 11.1% 76 of 182 Qwen3.8 2.4T A95B 7 Oct 2026
Artificial Analysis: Time to first answer tokenseconds, lower is better 54.3 115 of 128 Qwen3.8 2.4T A95B 7 Oct 2026
Artificial Analysis: τ-bench banking% of tasks, higher is better 49.1% 5 of 175 Qwen3.8 2.4T A95B 7 Oct 2026

Not shown: Not your codebase or tools.