Models / Qwen3.5-4B

Alibaba

Qwen3.5-4B

Qwen3.5-4B trails most of the field on what it has been measured on; it is not on sale through OpenRouter, so it is here for reference.

Compared with the models available today (it is not on sale itself), in the bottom quarter for writing.

Results as of 8 October 2026, from 30 results on 2 sources; prices checked 8 Oct 2026.

Price per million tokens
Not available through OpenRouter
Speed
24 tokens a second; on its test prompt the answer starts after 84.4 s, thinking time included
Artificial Analysis, on its usual API, at thinking reasoning; 20–24 across 2 settings
Intelligence Index
13.1
Artificial Analysis, at thinking reasoning; 10.8–13.1 across 2 settings
Spring Prompt overall
Not ranked yet
needs our benchmarks and two groups of results
Developer
Alibaba, based in China
Weights not published
Released
2 Mar 2026 (Artificial Analysis)

Where it stands among the models you could choose

Ranked only among models available today (those you can call through OpenRouter), in the groups people choose between: the same price band, similar intelligence, the same speed, developers based in the same place. Retired models are left out.

AmongIntelligence IndexSpring Prompt overallPrice, blendedSpeed
Models available today 122nd of 199Claude Opus 5.5 – – 65th of 65Trinity Large Thinking
Models with similar intelligence sets this group – – 18th of 18Trinity Large Thinking
Models writing under 50 tokens a second 7th of 10MiMo-V2.6-Pro – – sets this group
Models from developers based in China 56th of 87MiMo-V2.6-Pro – – 23rd of 23DeepSeek-V4.1-Flash

The name under each rank is the leader of that group; green is the top quarter of the group and red the bottom quarter, by rank. Price is shown as a percentile: the share of the group that costs less, so lower is cheaper. "Sets this group" marks the measure the group is defined by. Spring Prompt overall is our score out of 100 across every source (how it works). Price is the developer's list price (or the typical OpenRouter host where we haven't read one) for three input tokens to one output; intelligence and speed are from Artificial Analysis (data sourced from Artificial Analysis); "similar intelligence" means within 4 points on its Intelligence Index. Developer locations are where each company is based, not where a model is served. Groups under 3 models are not ranked.

How good is it, and for what?

Each use case ranks the available models its benchmarks measured. The bar shows its percentile in that field, best to the right. Open a row for the results behind it.

Product listingsTurning a sparse product feed and photos into listings that can go live Not measured
  • CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is Not measured
  • DeckBench: has not measured this model.
User surveysPlanning a user survey and reading its results without being misled Not measured
  • SurveyBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending Not measured
  • ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls 115th of 165 available
BenchmarkQwen3.5-4BRank among availableBest available
Artificial AnalysisArtificial Analysis: Terminal-Bench 2.1 · reported 25.8% 82nd of 116 Claude Fable 5.1 91.4%
Artificial AnalysisArtificial Analysis: τ-bench banking · reported 6.8% 87th of 114 Grok 4.6 50.7%
  • tau2-bench: has not measured this model.
  • Berkeley Function Calling Leaderboard (BFCL) V4: pending. Its latest results are dated 16 Dec 2025, before this model was released.
  • OpenHands Index: has not measured this model.
  • Microsoft STATE-Bench: has not measured this model.
  • Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects Not measured
  • APEX-Agents: has not measured this model.
  • GDP.pdf: has not measured this model.
  • Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge 125th of 199 available
BenchmarkQwen3.5-4BRank among availableBest available
Artificial AnalysisArtificial Analysis: Artificial Analysis Intelligence Index · reported 13.1 122nd of 199 Claude Opus 5.5 57.6
Artificial AnalysisArtificial Analysis: Humanity's Last Exam · reported 9.9% 134th of 198 Claude Opus 5.5 61.4%
Artificial AnalysisArtificial Analysis: GPQA Diamond · reported 77.1% 111th of 188 GPT-6 Astra 96.3%
WritingWhat people prefer in blind comparisons, and judged writing quality 146th of 189 available
BenchmarkQwen3.5-4BRank among availableBest available
UGI LeaderboardUGI: writing score · reported 30.8 111th of 139 Gemini 3.8 Flash 78.6
  • Arena (formerly LMArena): has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers Not measured
  • Vectara Hallucination Leaderboard: has not measured this model.
  • Arena (formerly LMArena): has not measured this model.
  • SimpleQA Verified (Epoch AI): has not measured this model.
Speed under pressureGood decisions against a real clock (fast chess) Not measured
  • BulletBench: has not measured this model.

Nearest alternatives

Against Qwen3.5-4B, Intelligence Index 13.1. "Similar intelligence" means within 4 points or better.

Against the alternatives

The models you would most likely weigh it against: the leaders of the groups above.

Scroll sideways to see every alternative.

Qwen3.5-4BTrinity Large Thinking
fastest with similar intelligence
Claude Opus 5.5
most intelligent available today
MiMo-V2.6-Pro
most intelligent from China
Price per million tokens – $0.39$8.00$0.54
Tokens a second 24 3479739
Intelligence Index 13.1 10.857.646.3
Spring Prompt overall – –85–
Rank among available models, by use case
Product listings – –6th–
Decks from an analysis – –5th–
User surveys – –2nd–
Marketing planning – –5th–
Agents and tool use 115th of 165 122nd2nd15th
Professional work – –3rd–
Reasoning and knowledge 125th of 199 126th1st10th
Writing 146th of 189 131st4th20th
Sticking to the facts – 142nd2nd8th
Speed under pressure – –––

Shaded figures are better than Qwen3.5-4B's. A dash means no result.

Where it has been measured

Every result

30 published results from 2 sources, each in the source's own units, with the configuration that produced it.

Artificial Analysis · reported by Artificial Analysis · 24 results
MeasureValueRankConfigurationDated
Artificial Analysis: Artificial Analysis Coding Indexindex score, higher is better 22.6 149 of 188 Qwen3.5-4B (thinking reasoning) 8 Oct 2026
Artificial Analysis: Artificial Analysis Coding Indexindex score, higher is better 20.3 157 of 188 Qwen3.5-4B (no reasoning) 8 Oct 2026
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 13.1 271 of 407 Qwen3.5-4B (thinking reasoning) 8 Oct 2026
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 10.8 306 of 407 Qwen3.5-4B (no reasoning) 8 Oct 2026
Artificial Analysis: GPQA Diamond% of questions, higher is better 77.1% 198 of 359 Qwen3.5-4B (thinking reasoning) 8 Oct 2026
Artificial Analysis: GPQA Diamond% of questions, higher is better 71.2% 241 of 359 Qwen3.5-4B (no reasoning) 8 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 9.9% 268 of 405 Qwen3.5-4B (thinking reasoning) 8 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 8.0% 284 of 405 Qwen3.5-4B (no reasoning) 8 Oct 2026
Artificial Analysis: IFBench% of instructions, higher is better 52.0% 132 of 288 Qwen3.5-4B (thinking reasoning) 8 Oct 2026
Artificial Analysis: IFBench% of instructions, higher is better 33.3% 247 of 288 Qwen3.5-4B (no reasoning) 8 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 63.0% 231 of 393 Qwen3.5-4B (thinking reasoning) 8 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 34.7% 313 of 393 Qwen3.5-4B (no reasoning) 8 Oct 2026
Artificial Analysis: Output speedtokens per second, higher is better 23.9 132 of 133 Qwen3.5-4B (thinking reasoning) 8 Oct 2026
Artificial Analysis: Output speedtokens per second, higher is better 19.8 133 of 133 Qwen3.5-4B (no reasoning) 8 Oct 2026
Artificial Analysis: Terminal-Bench 2.1% of tasks, higher is better 25.8% 141 of 187 Qwen3.5-4B (thinking reasoning) 8 Oct 2026
Artificial Analysis: Terminal-Bench 2.1% of tasks, higher is better 21.3% 142 of 187 Qwen3.5-4B (no reasoning) 8 Oct 2026
Artificial Analysis: Terminal-Bench Hard% of tasks, higher is better 18.2% 155 of 282 Qwen3.5-4B (thinking reasoning) 8 Oct 2026
Artificial Analysis: Terminal-Bench Hard% of tasks, higher is better 11.4% 194 of 282 Qwen3.5-4B (no reasoning) 8 Oct 2026
Artificial Analysis: Time to first answer tokenseconds, lower is better 84.4 126 of 133 Qwen3.5-4B (thinking reasoning) 8 Oct 2026
Artificial Analysis: Time to first answer tokenseconds, lower is better 0.71 16 of 133 Qwen3.5-4B (no reasoning) 8 Oct 2026
Artificial Analysis: τ-bench banking% of tasks, higher is better 6.8% 143 of 175 Qwen3.5-4B (thinking reasoning) 8 Oct 2026
Artificial Analysis: τ-bench banking% of tasks, higher is better 4.3% 166 of 175 Qwen3.5-4B (no reasoning) 8 Oct 2026
Artificial Analysis: τ²-bench telecom% of tasks, higher is better 92.1% 46 of 286 Qwen3.5-4B (thinking reasoning) 8 Oct 2026
Artificial Analysis: τ²-bench telecom% of tasks, higher is better 87.7% 64 of 286 Qwen3.5-4B (no reasoning) 8 Oct 2026

Not shown: Not your codebase or tools.

UGI Leaderboard · reported by UGI Leaderboard · 6 results
MeasureValueRankConfigurationDated
UGI: requested-length error% off the requested word count, lower is better 16.0% 135 of 379 Qwen3.5-4B (think-prefill reasoning) 15 Mar 2026
UGI: requested-length error% off the requested word count, lower is better 18.0% 158 of 379 Qwen3.5-4B (no reasoning) 14 Mar 2026
UGI: style adherencescore from 0 to 1, higher is better 0.31 331 of 379 Qwen3.5-4B (think-prefill reasoning) 15 Mar 2026
UGI: style adherencescore from 0 to 1, higher is better 0.32 298 of 379 Qwen3.5-4B (no reasoning) 14 Mar 2026
UGI: writing scorescore out of 100, higher is better 30.8 276 of 379 Qwen3.5-4B (think-prefill reasoning) 15 Mar 2026
UGI: writing scorescore out of 100, higher is better 29.7 286 of 379 Qwen3.5-4B (no reasoning) 14 Mar 2026

Not shown: Not other format limits such as character counts or bullet counts.