Models / Qwen3.8 Flash

Alibaba

Qwen3.8 Flash

Among models available today, in the bottom quarter for speed under pressure.

Price per million tokens
$0.15 in · $0.47 out
Typical host on OpenRouter, 7 Oct 2026 · compare prices
Speed
Not measured
Intelligence Index
Not measured
Spring Prompt overall
Not ranked yet
needs our benchmarks and two groups of results
Developer
Alibaba, based in China
Weights published
Released
Not recorded
1,000,000 tokens of context

Where it stands among the models you could choose

Ranked only among models available today (those you can call through OpenRouter), in the groups people choose between: the same price band, similar intelligence, the same speed, developers based in the same place, open weights. Retired models are left out.

AmongIntelligence IndexSpring Prompt overallPriceSpeed
Models available today – – 47th of 221Mistral Nemo –
Models from developers based in China – – 19th of 96Qwen3.5-9B –
Open-weights modelsincluding 1 with weights announced but not yet published – – 37th of 117Mistral Nemo –

The name under each rank is the leader of that group. Price is the developer's list price (or the typical OpenRouter host where we haven't read one) for three input tokens to one output; intelligence and speed are from Artificial Analysis (data sourced from Artificial Analysis); "similar intelligence" means within 4 points on its Intelligence Index. Developer locations are where each company is based, not where a model is served. Groups under 3 models are not ranked.

How good is it, and for what?

Each use case ranks the available models its benchmarks measured, then those priced under $0.25 per million tokens. The bar shows where it falls in that field, best to the right. Open a row for the results behind it.

Product listingsTurning a sparse product feed and photos into listings that can go live Not measured
  • CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is Not measured
  • DeckBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending Not measured
  • ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls Not measured
  • tau2-bench: has not measured this model.
  • Berkeley Function Calling Leaderboard (BFCL) V4: has not measured this model.
  • OpenHands Index: has not measured this model.
  • Microsoft STATE-Bench: has not measured this model.
  • Artificial Analysis: has not measured this model.
  • Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects 32nd of 51 available · 3rd of 3 at its price
BenchmarkQwen3.8 FlashRank among availableBest available
GDP.pdfGDP.pdf: professional document tasks · reported 16.6% 21st of 31 GPT-5.6 Sol 30.7%
  • APEX-Agents: has not measured this model.
  • Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge Not measured
  • Artificial Analysis: has not measured this model.
WritingWhat people prefer in blind comparisons, and judged writing quality Not measured
  • Arena (formerly LMArena): has not measured this model.
  • UGI Leaderboard: has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers Not measured
  • Vectara Hallucination Leaderboard: has not measured this model.
  • Arena (formerly LMArena): has not measured this model.
  • SimpleQA Verified (Epoch AI): has not measured this model.
Speed under pressureGood decisions against a real clock (fast chess) Too slow for the clock
  • BulletBench: too slow for Bullet 60s: lost its first 4 games on time, at a median 3.8 s a move; too slow for Lightning 10+1: lost its first 4 games on time, at a median 3.0 s a move.

Against the alternatives

The models you would most likely weigh it against: the leaders of the groups above.

Qwen3.8 FlashGPT-6.1 Sol
leads overall
Qwen3.8 Max (0902)
Alibaba's best other model
Claude Opus 5.5
near the top overall
GPT-6 Astra
near the top overall
Price per million tokens $0.23 $4.00$3.00$8.00$20.00
Tokens a second – 57399752
Intelligence Index – 51.845.457.652.7
Spring Prompt overall – 88468278
Rank among available models, by use case
Product listings – 2nd16th6th1st
Decks from an analysis – 3rd10th5th1st
Marketing planning – –14th5th1st
Agents and tool use – 3rd7th2nd4th
Professional work 32nd of 51 9th18th3rd5th
Reasoning and knowledge – 5th17th1st3rd
Writing – –22nd4th17th
Sticking to the facts – 2nd39th6th27th
Speed under pressure 22nd of 22 –––14th

Shaded figures are better than Qwen3.8 Flash's. A dash means no result.

Where it has been measured

Every result

1 published results from 1 source, each in the source's own units, with the configuration that produced it.

GDP.pdf · reported by GDP.pdf · 1 result
MeasureValueRankConfigurationDated
GDP.pdf: professional document tasks% of rubric, higher is better 16.6% 30 of 49 Qwen3.8 Flash (xhigh) 1 Oct 2026

Not shown: Not your documents; 100 tasks, so small differences are noise.