Models / Devstral Small 1.1

Mistral AI

Devstral Small 1.1

Early evidence, 7 results from 1 source, on 1 use case. Devstral Small 1.1 trails most of the field on what it has been measured on; it is not on sale through OpenRouter, so it is here for reference.

Compared with the models available today (it is not on sale itself), in the bottom quarter for reasoning and knowledge.

Results as of 8 October 2026, from 7 results on 1 source; prices checked 8 Oct 2026.

Price per million tokens
Not available through OpenRouter
Speed
No speed yet
Artificial Analysis lists it but has not published a speed for it yet
Intelligence Index
7.6
Artificial Analysis
Spring Prompt overall
Not ranked yet
needs our benchmarks and two groups of results
Developer
Mistral AI, based in France
Weights not published
Released
10 Jul 2025 (Artificial Analysis)

Where it stands among the models you could choose

Ranked only among models available today (those you can call through OpenRouter), in the groups people choose between: the same price band, similar intelligence, the same speed, developers based in the same place. Retired models are left out.

AmongIntelligence IndexSpring Prompt overallPrice, blendedSpeed
Models available today 171st of 199Claude Opus 5.5 – – –
Models from developers based in the EUall Mistral AI's 9th of 16Mistral Large 4 – – –

The name under each rank is the leader of that group; green is the top quarter of the group and red the bottom quarter, by rank. Price is shown as a percentile: the share of the group that costs less, so lower is cheaper. "Sets this group" marks the measure the group is defined by. Spring Prompt overall is our score out of 100 across every source (how it works). Price is the developer's list price (or the typical OpenRouter host where we haven't read one) for three input tokens to one output; intelligence and speed are from Artificial Analysis (data sourced from Artificial Analysis); "similar intelligence" means within 4 points on its Intelligence Index. Developer locations are where each company is based, not where a model is served. Groups under 3 models are not ranked.

How good is it, and for what?

Each use case ranks the available models its benchmarks measured. The bar shows its percentile in that field, best to the right. Open a row for the results behind it.

Product listingsTurning a sparse product feed and photos into listings that can go live Not measured
  • CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is Not measured
  • DeckBench: has not measured this model.
User surveysPlanning a user survey and reading its results without being misled Not measured
  • SurveyBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending Not measured
  • ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls Not measured
  • tau2-bench: has not measured this model.
  • Berkeley Function Calling Leaderboard (BFCL) V4: has not measured this model.
  • OpenHands Index: has not measured this model.
  • Microsoft STATE-Bench: has not measured this model.
  • Artificial Analysis: has not measured this model.
  • Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects Not measured
  • APEX-Agents: has not measured this model.
  • GDP.pdf: has not measured this model.
  • Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge 188th of 199 available
BenchmarkDevstral Small 1.1Rank among availableBest available
Artificial AnalysisArtificial Analysis: Artificial Analysis Intelligence Index · reported 7.6 171st of 199 Claude Opus 5.5 57.6
Artificial AnalysisArtificial Analysis: Humanity's Last Exam · reported 3.8% 182nd of 198 Claude Opus 5.5 61.4%
Artificial AnalysisArtificial Analysis: GPQA Diamond · reported 41.4% 179th of 188 GPT-6 Astra 96.3%
WritingWhat people prefer in blind comparisons, and judged writing quality Not measured
  • Arena (formerly LMArena): has not measured this model.
  • UGI Leaderboard: has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers Not measured
  • Vectara Hallucination Leaderboard: has not measured this model.
  • Arena (formerly LMArena): has not measured this model.
  • SimpleQA Verified (Epoch AI): has not measured this model.
Speed under pressureGood decisions against a real clock (fast chess) Not measured
  • BulletBench: has not measured this model.

Nearest alternatives

  • Mistral AI's newer, at least as capable modelMistral Large 4$2.06 blended · Intelligence Index 38.4
  • Open weights at similar intelligenceMistral Small 3$0.06 blended · Intelligence Index 6.7

Against Devstral Small 1.1, Intelligence Index 7.6. "Similar intelligence" means within 4 points or better.

Against the alternatives

The models you would most likely weigh it against: the leaders of the groups above.

Scroll sideways to see every alternative.

Devstral Small 1.1Claude Opus 5.5
most intelligent available today
Mistral Large 4
most intelligent from the EU
GPT-6.1 Sol
near the top overall
GPT-6 Astra
near the top overall
Price per million tokens – $8.00$2.06$4.00$20.00
Tokens a second – 971065552
Intelligence Index 7.6 57.638.451.852.7
Spring Prompt overall – 85378179
Rank among available models, by use case
Product listings – 6th18th2nd1st
Decks from an analysis – 5th19th3rd1st
User surveys – 2nd4th1st–
Marketing planning – 5th10th–1st
Agents and tool use – 2nd21st3rd4th
Professional work – 3rd–9th5th
Reasoning and knowledge 188th of 199 1st44th5th3rd
Writing – 4th–15th19th
Sticking to the facts – 2nd–18th24th
Speed under pressure – –13th–14th

Shaded figures are better than Devstral Small 1.1's. A dash means no result.

Where it has been measured

Every result

7 published results from 1 source, each in the source's own units, with the configuration that produced it.

Artificial Analysis · reported by Artificial Analysis · 7 results
MeasureValueRankConfigurationDated
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 7.6 358 of 407 Devstral Small 1.1 8 Oct 2026
Artificial Analysis: GPQA Diamond% of questions, higher is better 41.4% 343 of 359 Devstral Small 1.1 8 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 3.8% 375 of 405 Devstral Small 1.1 8 Oct 2026
Artificial Analysis: IFBench% of instructions, higher is better 34.6% 239 of 288 Devstral Small 1.1 8 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 18.7% 354 of 393 Devstral Small 1.1 8 Oct 2026
Artificial Analysis: Terminal-Bench Hard% of tasks, higher is better 6.1% 224 of 282 Devstral Small 1.1 8 Oct 2026
Artificial Analysis: τ²-bench telecom% of tasks, higher is better 28.4% 220 of 286 Devstral Small 1.1 8 Oct 2026

Not shown: Not business work, and a blend: read the parts for any one task.