Models / Qwen3.5-0.8B
Qwen3.5-0.8B
Qwen3.5-0.8B trails most of the field on what it has been measured on; it is not on sale through OpenRouter, so it is here for reference.
Compared with the models available today (it is not on sale itself), in the bottom quarter for agents and tool use, reasoning and knowledge, and writing.
Results as of 8 October 2026, from 21 results on 2 sources; prices checked 8 Oct 2026.
- Price per million tokens
- Not available through OpenRouter
- Speed
- No speed yet
Artificial Analysis lists it but has not published a speed for it yet - Intelligence Index
- 6.1
Artificial Analysis, at thinking reasoning; 5.4–6.1 across 2 settings - Spring Prompt overall
- Not ranked yet
needs our benchmarks and two groups of results - Developer
- Alibaba, based in China
Weights not published - Released
- 2 Mar 2026 (Artificial Analysis)
Where it stands among the models you could choose
Ranked only among models available today (those you can call through OpenRouter), in the groups people choose between: the same price band, similar intelligence, the same speed, developers based in the same place. Retired models are left out.
| Among | Intelligence Index | Spring Prompt overall | Price, blended | Speed |
|---|---|---|---|---|
| Models available today | 187th of 199Claude Opus 5.5 | – | – | – |
| Models from developers based in China | 87th of 87MiMo-V2.6-Pro | – | – | – |
The name under each rank is the leader of that group; green is the top quarter of the group and red the bottom quarter, by rank. Price is shown as a percentile: the share of the group that costs less, so lower is cheaper. "Sets this group" marks the measure the group is defined by. Spring Prompt overall is our score out of 100 across every source (how it works). Price is the developer's list price (or the typical OpenRouter host where we haven't read one) for three input tokens to one output; intelligence and speed are from Artificial Analysis (data sourced from Artificial Analysis); "similar intelligence" means within 4 points on its Intelligence Index. Developer locations are where each company is based, not where a model is served. Groups under 3 models are not ranked.
How good is it, and for what?
Each use case ranks the available models its benchmarks measured. The bar shows its percentile in that field, best to the right. Open a row for the results behind it.
Product listingsTurning a sparse product feed and photos into listings that can go live
- CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is
- DeckBench: has not measured this model.
User surveysPlanning a user survey and reading its results without being misled
- SurveyBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending
- ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls
| Benchmark | Qwen3.5-0.8B | Rank among available | Best available |
|---|---|---|---|
| Artificial AnalysisArtificial Analysis: Terminal-Bench 2.1 · reported | 0.4% | 113th of 116 | Claude Fable 5.1 91.4% |
- tau2-bench: has not measured this model.
- Berkeley Function Calling Leaderboard (BFCL) V4: pending. Its latest results are dated 16 Dec 2025, before this model was released.
- OpenHands Index: has not measured this model.
- Microsoft STATE-Bench: has not measured this model.
- Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects
- APEX-Agents: has not measured this model.
- GDP.pdf: has not measured this model.
- Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge
| Benchmark | Qwen3.5-0.8B | Rank among available | Best available |
|---|---|---|---|
| Artificial AnalysisArtificial Analysis: Artificial Analysis Intelligence Index · reported | 6.1 | 187th of 199 | Claude Opus 5.5 57.6 |
| Artificial AnalysisArtificial Analysis: Humanity's Last Exam · reported | 5.1% | 159th of 198 | Claude Opus 5.5 61.4% |
| Artificial AnalysisArtificial Analysis: GPQA Diamond · reported | 23.6% | 187th of 188 | GPT-6 Astra 96.3% |
WritingWhat people prefer in blind comparisons, and judged writing quality
| Benchmark | Qwen3.5-0.8B | Rank among available | Best available |
|---|---|---|---|
| UGI LeaderboardUGI: writing score · reported | 23.0 | 132nd of 139 | Gemini 3.8 Flash 78.6 |
- Arena (formerly LMArena): has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers
- Vectara Hallucination Leaderboard: has not measured this model.
- Arena (formerly LMArena): has not measured this model.
- SimpleQA Verified (Epoch AI): has not measured this model.
Speed under pressureGood decisions against a real clock (fast chess)
- BulletBench: has not measured this model.
Nearest alternatives
- Alibaba's newer, at least as capable modelQwen3.8-Max (0902)$3.00 blended · Intelligence Index 45.4
- Open weights at similar intelligenceMistral Small 3$0.06 blended · Intelligence Index 6.7Compare with Qwen3.5-0.8B →
Against Qwen3.5-0.8B, Intelligence Index 6.1. "Similar intelligence" means within 4 points or better.
Against the alternatives
The models you would most likely weigh it against: the leaders of the groups above.
Scroll sideways to see every alternative.
| Qwen3.5-0.8B | Claude Opus 5.5 most intelligent available today | MiMo-V2.6-Pro most intelligent from China | Qwen3.8-Max (0902) Alibaba's best other model | GPT-6.1 Sol near the top overall | |
|---|---|---|---|---|---|
| Price per million tokens | – | $8.00 | $0.54 | $3.00 | $4.00 |
| Tokens a second | – | 97 | 39 | 37 | 55 |
| Intelligence Index | 6.1 | 57.6 | 46.3 | 45.4 | 51.8 |
| Spring Prompt overall | – | 85 | – | 48 | 81 |
| Rank among available models, by use case | |||||
| Product listings | – | 6th | – | 17th | 2nd |
| Decks from an analysis | – | 5th | – | 11th | 3rd |
| User surveys | – | 2nd | – | – | 1st |
| Marketing planning | – | 5th | – | 14th | – |
| Agents and tool use | 163rd of 165 | 2nd | 15th | 7th | 3rd |
| Professional work | – | 3rd | – | 18th | 9th |
| Reasoning and knowledge | 184th of 199 | 1st | 10th | 17th | 5th |
| Writing | 180th of 189 | 4th | 20th | 16th | 15th |
| Sticking to the facts | – | 2nd | 8th | 38th | 18th |
| Speed under pressure | – | – | – | – | – |
| Side by side → | Side by side → | ||||
Shaded figures are better than Qwen3.5-0.8B's. A dash means no result.
Where it has been measured
- Artificial Analysis18 results
- UGI Leaderboard3 results
- Berkeley Function Calling Leaderboard (BFCL) V4Pending: latest results 16 Dec 2025
- BulletBenchNot measured
- CatalogBenchNot measured
- DeckBenchNot measured
- ROASBenchNot measured
- SurveyBenchNot measured
- APEX-AgentsNot measured
- Arena (formerly LMArena)Not measured
- GDP.pdfNot measured
- Microsoft STATE-BenchNot measured
- OpenHands IndexNot measured
- Remote Labor IndexNot measured
- SimpleQA Verified (Epoch AI)Not measured
- tau2-benchNot measured
- Vectara Hallucination LeaderboardNot measured
- Vending-Bench 2Not measured
Every result
21 published results from 2 sources, each in the source's own units, with the configuration that produced it.
Artificial Analysis · reported by Artificial Analysis · 18 results
| Measure | Value | Rank | Configuration | Dated |
|---|---|---|---|---|
| Artificial Analysis: Artificial Analysis Coding Indexindex score, higher is better | 0.0 | 188 of 188 | Qwen3.5-0.8B (thinking reasoning) | 8 Oct 2026 |
| Artificial Analysis: Artificial Analysis Coding Indexindex score, higher is better | 1.2 | 187 of 188 | Qwen3.5-0.8B (no reasoning) | 8 Oct 2026 |
| Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better | 6.1 | 390 of 407 | Qwen3.5-0.8B (thinking reasoning) | 8 Oct 2026 |
| Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better | 5.4 | 401 of 407 | Qwen3.5-0.8B (no reasoning) | 8 Oct 2026 |
| Artificial Analysis: GPQA Diamond% of questions, higher is better | 11.1% | 359 of 359 | Qwen3.5-0.8B (thinking reasoning) | 8 Oct 2026 |
| Artificial Analysis: GPQA Diamond% of questions, higher is better | 23.6% | 356 of 359 | Qwen3.5-0.8B (no reasoning) | 8 Oct 2026 |
| Artificial Analysis: Humanity's Last Exam% of questions, higher is better | 1.1% | 405 of 405 | Qwen3.5-0.8B (thinking reasoning) | 8 Oct 2026 |
| Artificial Analysis: Humanity's Last Exam% of questions, higher is better | 5.1% | 328 of 405 | Qwen3.5-0.8B (no reasoning) | 8 Oct 2026 |
| Artificial Analysis: IFBench% of instructions, higher is better | 21.5% | 287 of 288 | Qwen3.5-0.8B (thinking reasoning) | 8 Oct 2026 |
| Artificial Analysis: IFBench% of instructions, higher is better | 21.6% | 286 of 288 | Qwen3.5-0.8B (no reasoning) | 8 Oct 2026 |
| Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better | 9.0% | 368 of 393 | Qwen3.5-0.8B (thinking reasoning) | 8 Oct 2026 |
| Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better | 8.0% | 370 of 393 | Qwen3.5-0.8B (no reasoning) | 8 Oct 2026 |
| Artificial Analysis: Terminal-Bench 2.1% of tasks, higher is better | 0.0% | 184 of 187 | Qwen3.5-0.8B (thinking reasoning) | 8 Oct 2026 |
| Artificial Analysis: Terminal-Bench 2.1% of tasks, higher is better | 0.4% | 181 of 187 | Qwen3.5-0.8B (no reasoning) | 8 Oct 2026 |
| Artificial Analysis: Terminal-Bench Hard% of tasks, higher is better | 0.0% | 276 of 282 | Qwen3.5-0.8B (thinking reasoning) | 8 Oct 2026 |
| Artificial Analysis: Terminal-Bench Hard% of tasks, higher is better | 0.0% | 276 of 282 | Qwen3.5-0.8B (no reasoning) | 8 Oct 2026 |
| Artificial Analysis: τ²-bench telecom% of tasks, higher is better | 47.7% | 165 of 286 | Qwen3.5-0.8B (thinking reasoning) | 8 Oct 2026 |
| Artificial Analysis: τ²-bench telecom% of tasks, higher is better | 65.2% | 141 of 286 | Qwen3.5-0.8B (no reasoning) | 8 Oct 2026 |
Not shown: Not your codebase or tools.
UGI Leaderboard · reported by UGI Leaderboard · 3 results
| Measure | Value | Rank | Configuration | Dated |
|---|---|---|---|---|
| UGI: requested-length error% off the requested word count, lower is better | 24.0% | 225 of 379 | Qwen3.5-0.8B (no reasoning) | 14 Mar 2026 |
| UGI: style adherencescore from 0 to 1, higher is better | 0.32 | 306 of 379 | Qwen3.5-0.8B (no reasoning) | 14 Mar 2026 |
| UGI: writing scorescore out of 100, higher is better | 23.0 | 328 of 379 | Qwen3.5-0.8B (no reasoning) | 14 Mar 2026 |
Not shown: Not other format limits such as character counts or bullet counts.