Models / Devstral 2 2512

Mistral

Devstral 2 2512

56th most intelligent of the 68 models priced $0.25–1 per million tokens. Among models available today, in the bottom quarter for reasoning and knowledge.

Price per million tokens
$0.40 in · $2.00 out
Typical host on OpenRouter, 7 Oct 2026 · compare prices
Speed
Not measured
Intelligence Index
8.6
Artificial Analysis
Spring Prompt overall
Not ranked yet
needs our benchmarks and two groups of results
Developer
Mistral, based in France
Weights published
Released
Not recorded
262,144 tokens of context

Where it stands among the models you could choose

Ranked only among models available today (those you can call through OpenRouter), in the groups people choose between: the same price band, similar intelligence, the same speed, developers based in the same place, open weights. Retired models are left out.

AmongIntelligence IndexSpring Prompt overallPriceSpeed
Models available today 153rd of 195Claude Opus 5.5 – 112th of 221Mistral Nemo –
Models priced $0.25–1 per million tokens 56th of 68MiMo-V2.6-Pro – the group –
Models with similar intelligence the group – 57th of 69Llama 3.1 8B Instruct –
Models from developers based in the EUall Mistral's 7th of 15Mistral Large 4 – 10th of 16Mistral Nemo –
Open-weights modelsincluding 1 with weights announced but not yet published 76th of 108MiMo-V2.6-Pro – 85th of 117Mistral Nemo –

The name under each rank is the leader of that group. Price is the developer's list price (or the typical OpenRouter host where we haven't read one) for three input tokens to one output; intelligence and speed are from Artificial Analysis (data sourced from Artificial Analysis); "similar intelligence" means within 4 points on its Intelligence Index. Developer locations are where each company is based, not where a model is served. Groups under 3 models are not ranked.

How good is it, and for what?

Each use case ranks the available models its benchmarks measured, then those priced $0.25–1 per million tokens. The bar shows where it falls in that field, best to the right. Open a row for the results behind it.

Product listingsTurning a sparse product feed and photos into listings that can go live Not measured
  • CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is Not measured
  • DeckBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending Not measured
  • ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls 104th of 161 available · 25th of 53 at its price
BenchmarkDevstral 2 2512Rank among availableBest available
Artificial AnalysisArtificial Analysis: Terminal-Bench 4.0 · reported 0.0% 58th of 102 Claude Sonnet 5.5 63.6%
Artificial AnalysisArtificial Analysis: Terminal-Bench 2.1 · reported 30.3% 74th of 113 Claude Fable 5.1 91.4%
Artificial AnalysisArtificial Analysis: τ-bench banking · reported 10.5% 67th of 110 Grok 4.6 50.7%
  • tau2-bench: has not measured this model.
  • Berkeley Function Calling Leaderboard (BFCL) V4: has not measured this model.
  • OpenHands Index: has not measured this model.
  • Microsoft STATE-Bench: has not measured this model.
  • Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects Not measured
  • APEX-Agents: has not measured this model.
  • GDP.pdf: has not measured this model.
  • Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge 161st of 195 available · 59th of 68 at its price
BenchmarkDevstral 2 2512Rank among availableBest available
Artificial AnalysisArtificial Analysis: Artificial Analysis Intelligence Index · reported 8.60 153rd of 195 Claude Opus 5.5 57.6
Artificial AnalysisArtificial Analysis: Humanity's Last Exam · reported 3.6% 185th of 194 Claude Opus 5.5 61.4%
Artificial AnalysisArtificial Analysis: GPQA Diamond · reported 59.4% 148th of 185 GPT-6 Astra 96.3%
WritingWhat people prefer in blind comparisons, and judged writing quality Not measured
  • Arena (formerly LMArena): has not measured this model.
  • UGI Leaderboard: has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers Not measured
  • Vectara Hallucination Leaderboard: has not measured this model.
  • Arena (formerly LMArena): has not measured this model.
  • SimpleQA Verified (Epoch AI): has not measured this model.
Speed under pressureGood decisions against a real clock (fast chess) Not measured
  • BulletBench: has not measured this model.

Against the alternatives

The models you would most likely weigh it against: the leaders of the groups above.

Devstral 2 2512MiMo-V2.6-Pro
most intelligent at $0.25–1 per million tokens
Llama 3.1 8B Instruct
cheapest with similar intelligence
Claude Opus 5.5
most intelligent available today
Mistral Large 4
most intelligent from the EU
Price per million tokens $0.80 $0.54$0.06$8.00$1.03
Tokens a second – 40–97106
Intelligence Index 8.6 46.36.957.638.4
Spring Prompt overall – ––8231
Rank among available models, by use case
Product listings – ––6th17th
Decks from an analysis – ––5th18th
Marketing planning – ––5th10th
Agents and tool use 104th of 161 15th154th2nd20th
Professional work – ––3rd–
Reasoning and knowledge 161st of 195 10th173rd1st43rd
Writing – 18th178th4th–
Sticking to the facts – 9th–6th–
Speed under pressure – –––13th

Shaded figures are better than Devstral 2 2512's. A dash means no result.

Where it has been measured

Every result

12 published results from 1 source, each in the source's own units, with the configuration that produced it.

Artificial Analysis · reported by Artificial Analysis · 12 results
MeasureValueRankConfigurationDated
Artificial Analysis: Artificial Analysis Coding Indexindex score, higher is better 31.3 134 of 188 Devstral 2 2512 7 Oct 2026
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 8.60 333 of 402 Devstral 2 2512 7 Oct 2026
Artificial Analysis: GPQA Diamond% of questions, higher is better 59.4% 294 of 359 Devstral 2 2512 7 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 3.6% 382 of 400 Devstral 2 2512 7 Oct 2026
Artificial Analysis: IFBench% of instructions, higher is better 38.1% 217 of 288 Devstral 2 2512 7 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 32.3% 314 of 388 Devstral 2 2512 7 Oct 2026
Artificial Analysis: SciCode% of problems, higher is better 32.8% 168 of 186 Devstral 2 2512 7 Oct 2026
Artificial Analysis: Terminal-Bench 2.1% of tasks, higher is better 30.3% 132 of 187 Devstral 2 2512 7 Oct 2026
Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better 0.0% 132 of 182 Devstral 2 2512 7 Oct 2026
Artificial Analysis: Terminal-Bench Hard% of tasks, higher is better 18.9% 150 of 282 Devstral 2 2512 7 Oct 2026
Artificial Analysis: τ-bench banking% of tasks, higher is better 10.5% 121 of 175 Devstral 2 2512 7 Oct 2026
Artificial Analysis: τ²-bench telecom% of tasks, higher is better 24.9% 236 of 286 Devstral 2 2512 7 Oct 2026

Not shown: Not your codebase or tools.