Compare / Mistral Large 4 vs Mistral Medium 3.5

Mistral Large 4 vs Mistral Medium 3.5

Which is cheaper, which is more capable, and which does better at each kind of work, from 72 results on 5 sources that measured both models.

Mistral Large 4
Mistral · profile
Mistral Medium 3.5
Mistral · profile
Results compared
72
Shared sources
5

Mistral Large 4 or Mistral Medium 3.5? The short answer

  • Mistral Large 4 is cheaper: $1.03 against $3.00 per million tokens, 2.9 times less.
  • Mistral Large 4 scores higher on the Artificial Analysis Intelligence Index: 38.4 against 14.2.
  • Mistral Medium 3.5 writes faster: 163 tokens a second against 106.
  • On our business benchmarks, Mistral Large 4 ranks higher for marketing planning, agents and tool use and reasoning and knowledge.
  • On our business benchmarks, Mistral Medium 3.5 ranks higher for product listings, decks from an analysis and speed under pressure.

At a glance

Mistral Large 4Mistral Medium 3.5
Price per million tokens$0.68 in · $2.09 out$1.50 in · $7.50 out
Blended (3 in : 1 out)$1.03$3.00
Cheapest host, blended–$3.00 (Mistral)
Intelligence Index38.414.2
Spring Prompt overall3121
Output speed, tokens a second106163
Context window524,288 tokens262,144 tokens
WeightsClosedClosed
Developer based inFranceFrance

Prices are the developer's list price where we have read it, otherwise the typical host on OpenRouter; see every model's price. Intelligence and speed from Artificial Analysis.

Which is better for what

Each model's rank among models available today, for each kind of work, from our own benchmarks and licensed sources. The better rank is in green.

Use caseMistral Large 4Mistral Medium 3.5
Product listingsTurning a sparse product feed and photos into listings that can go live 17 of 19 13 of 19
Decks from an analysisTurning a finished analysis into a deck you could present as it is 18 of 18 14 of 18
Marketing planningPlanning a year of ad spend without overspending 10 of 18 18 of 18
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls 20 of 161 87 of 161
Reasoning and knowledgeHard questions across science, maths and general knowledge 43 of 195 119 of 195
WritingWhat people prefer in blind comparisons, and judged writing quality not measured 85 of 187
Sticking to the factsSummarising without inventing things, and factual answers not measured 99 of 153
Speed under pressureGood decisions against a real clock (fast chess) 13 of 21 8 of 21

Their slides, side by side

The same brief for both: Manchester store investment review for Hearthside Coffee, an invented company. First slide of each deck from DeckBench.

First slide of Mistral Large 4's deck for Hearthside Coffee
Mistral Large 4 · deck rating 515 · needs work
First slide of Mistral Medium 3.5's deck for Hearthside Coffee
Mistral Medium 3.5 · deck rating 649 · needs work

Every slide, side by side →

From our research

Every shared result

BenchmarkMetricMistral Large 4Mistral Medium 3.5
BulletBenchmeasured by usCost per gameUS dollars $0.0082$0.0261
BulletBenchmeasured by usGames lost on time% of games 0.0%0.0%
BulletBenchmeasured by usInvalid moves% of moves 0.6%0.4%
BulletBenchmeasured by usLadder Eloladder Elo 555555
BulletBenchmeasured by usMedian move timemilliseconds 0.8 s0.5 s
BulletBenchmeasured by usCost per gameUS dollars $0.0117$0.0297
BulletBenchmeasured by usGames lost on time% of games 58.3%0.0%
BulletBenchmeasured by usInvalid moves% of moves 1.8%0.0%
BulletBenchmeasured by usLadder Eloladder Elo 249529
BulletBenchmeasured by usMedian move timemilliseconds 0.8 s0.5 s
CatalogBenchmeasured by usChannel rules broken% of products 1.2%1.2%
CatalogBenchmeasured by usClaims to check% of products 3.3%25.6%
CatalogBenchmeasured by usContent quality% of checks 95.9%94.1%
CatalogBenchmeasured by usFailed outputs% of products 0.0%0.0%
CatalogBenchmeasured by usMissing UK information% of products 0.0%0.0%
CatalogBenchmeasured by usNot findable% of products 3.0%10.1%
CatalogBenchmeasured by usPublish-ready listings% of products 10.7%17.3%
CatalogBenchmeasured by usReliably publish-ready% of products 1.8%10.7%
CatalogBenchmeasured by usUnsupported claimsclaims per product 3.693.55
CatalogBenchmeasured by usUnsupported claims% of products 80.4%76.8%
CatalogBenchmeasured by usWrong attributes% of products 17.3%36.3%
CatalogBenchmeasured by usWrong category or variant% of products 1.8%13.1%
CatalogBenchmeasured by usChannel compliance% of products 100.0%100.0%
CatalogBenchmeasured by usChannel rules broken% of products 3.0%0.0%
CatalogBenchmeasured by usClaims to check% of products 0.7%3.6%
CatalogBenchmeasured by usConflicts caught% of conflicts 100.0%90.9%
CatalogBenchmeasured by usContent quality% of checks 96.5%94.1%
CatalogBenchmeasured by usCost per productUS dollars $0.0033$0.0080
CatalogBenchmeasured by usDecision accuracy% of decisions 93.1%92.6%
CatalogBenchmeasured by usFailed outputs% of products 0.0%0.0%
CatalogBenchmeasured by usField accuracy% of missing fields 94.4%91.7%
CatalogBenchmeasured by usInvented values% of filled values 3.5%0.0%
CatalogBenchmeasured by usMissing UK information% of products 0.0%0.0%
CatalogBenchmeasured by usNot findable% of products 4.2%14.3%
CatalogBenchmeasured by usPublish-ready listings% of products 39.9%29.8%
CatalogBenchmeasured by usReliably publish-ready% of products 10.7%17.9%
CatalogBenchmeasured by usUnsupported claims% of products 33.3%49.4%
CatalogBenchmeasured by usUnsupported claimsclaims per product 0.911.74
CatalogBenchmeasured by usWrong attributes% of products 19.6%30.9%
CatalogBenchmeasured by usWrong category or variant% of products 1.2%11.3%
DeckBenchmeasured by usAccurate decks% of tasks 0.0%0.0%
DeckBenchmeasured by usCaveat dropped% of tasks 16.7%0.0%
DeckBenchmeasured by usClean layout% of tasks 33.3%66.7%
DeckBenchmeasured by usCost per deckUS dollars $0.0148$0.0448
DeckBenchmeasured by usDeck ratingrating 515649
DeckBenchmeasured by usDesign quality% of the maximum 30.5%43.6%
DeckBenchmeasured by usDraft figure quoted% of tasks 16.7%50.0%
DeckBenchmeasured by usFindings missing% of tasks 0.0%0.0%
DeckBenchmeasured by usHead-to-head win rate% of comparisons 10.7%18.6%
DeckBenchmeasured by usLayout defects% of tasks 66.7%33.3%
DeckBenchmeasured by usMisleading metric used% of tasks 0.0%0.0%
DeckBenchmeasured by usPresentable decks% of tasks 0.0%0.0%
DeckBenchmeasured by usRecommendation late or wrong% of tasks 83.3%50.0%
DeckBenchmeasured by usSlides needing work% of tasks 100.0%50.0%
DeckBenchmeasured by usUnsupported claims% of tasks 50.0%33.3%
DeckBenchmeasured by usUnsupported numbers% of tasks 33.3%0.0%
ROASBenchmeasured by usAudience scorescore out of 100 55.828.5
ROASBenchmeasured by usBusiness scorescore out of 100 34.23.4
ROASBenchmeasured by usConsistency scorescore out of 100 16.78.7
ROASBenchmeasured by usContribution profitsimulated US dollars $533,401−$159,882
ROASBenchmeasured by usCost of a runUS dollars $0.0657$0.15
ROASBenchmeasured by usMonths over budgetmonths of 12 03
ROASBenchmeasured by usOverall scorescore out of 100 37.015.2
ROASBenchmeasured by usPlanning scorescore out of 100 55.055.2
ROASBenchmeasured by usReturn on ad spendprofit per $1 spent $0.45−$0.17
Artificial AnalysisreportedArtificial Analysis Intelligence Indexindex score 38.414.2
Artificial AnalysisreportedHumanity's Last Exam% of questions 35.0%13.8%
Artificial AnalysisreportedLong-context reasoning (AA-LCR)% of questions 81.3%69.3%
Artificial AnalysisreportedOutput speedtokens per second 106163
Artificial AnalysisreportedSciCode% of problems 54.2%40.2%
Artificial AnalysisreportedTerminal-Bench 4.0% of tasks 26.8%0.0%
Artificial AnalysisreportedTime to first answer tokenseconds 19.812.9

Bold green marks the better value on that metric. Where a source reports ranges that overlap, the difference may not be meaningful; see the benchmark page for ranges.

Quick answers

Is Mistral Large 4 better than Mistral Medium 3.5?

It depends on the task. Among models available today, on our benchmarks and licensed sources, Mistral Large 4 ranks higher for marketing planning, agents and tool use and reasoning and knowledge; Mistral Medium 3.5 ranks higher for product listings, decks from an analysis and speed under pressure.

Which is cheaper, Mistral Large 4 or Mistral Medium 3.5?

Mistral Large 4 costs $1.03 per million tokens (three input to one output), against $3.00 for Mistral Medium 3.5.

Which is faster, Mistral Large 4 or Mistral Medium 3.5?

Mistral Medium 3.5 writes about 163 tokens a second on its usual API, against 106 for Mistral Large 4 (Artificial Analysis).

Run Mistral Large 4 and Mistral Medium 3.5 on your own prompt

Benchmarks aren't your data. Try both side by side in the playground, with the cost of every answer.