Compare / Mistral Medium 3.5 vs GPT-6.1 Sol

Mistral Medium 3.5 vs GPT-6.1 Sol

Which is cheaper, which is more capable, and which does better at each kind of work, from 53 results on 3 sources that measured both models.

Mistral Medium 3.5
Mistral · profile
GPT-6.1 Sol
OpenAI · profile
Results compared
53
Shared sources
3

Mistral Medium 3.5 or GPT-6.1 Sol? The short answer

  • Mistral Medium 3.5 is cheaper: $3.00 against $4.00 per million tokens, 25% less.
  • GPT-6.1 Sol scores higher on the Artificial Analysis Intelligence Index: 51.8 against 14.2.
  • Mistral Medium 3.5 writes faster: 163 tokens a second against 57.
  • On our business benchmarks, GPT-6.1 Sol ranks higher for product listings, decks from an analysis, agents and tool use, reasoning and knowledge and sticking to the facts.

At a glance

Mistral Medium 3.5GPT-6.1 Sol
Price per million tokens$1.50 in · $7.50 out$2.00 in · $10.00 out
Blended (3 in : 1 out)$3.00$4.00
Cheapest host, blended$3.00 (Mistral)$2.00 (OpenAI)
Intelligence Index14.251.8
Spring Prompt overall2188
Output speed, tokens a second16357
Context window262,144 tokens1,050,000 tokens
WeightsClosedClosed
Developer based inFrancethe US

Prices are the developer's list price where we have read it, otherwise the typical host on OpenRouter; see every model's price. Intelligence and speed from Artificial Analysis.

Which is better for what

Each model's rank among models available today, for each kind of work, from our own benchmarks and licensed sources. The better rank is in green.

Use caseMistral Medium 3.5GPT-6.1 Sol
Product listingsTurning a sparse product feed and photos into listings that can go live 13 of 19 2 of 19
Decks from an analysisTurning a finished analysis into a deck you could present as it is 14 of 18 3 of 18
Marketing planningPlanning a year of ad spend without overspending 18 of 18 not measured
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls 87 of 161 3 of 161
Professional workReal tasks from banking, consulting and law, business documents and freelance projects not measured 9 of 51
Reasoning and knowledgeHard questions across science, maths and general knowledge 119 of 195 5 of 195
WritingWhat people prefer in blind comparisons, and judged writing quality 85 of 187 not measured
Sticking to the factsSummarising without inventing things, and factual answers 99 of 153 2 of 153
Speed under pressureGood decisions against a real clock (fast chess) 8 of 21 not measured

Their slides, side by side

The same brief for both: Manchester store investment review for Hearthside Coffee, an invented company. First slide of each deck from DeckBench.

First slide of Mistral Medium 3.5's deck for Hearthside Coffee
Mistral Medium 3.5 · deck rating 649 · needs work
First slide of GPT-6.1 Sol's deck for Hearthside Coffee
GPT-6.1 Sol · deck rating 1,412 · presentable as it is

Every slide, side by side →

From our research

Every shared result

BenchmarkMetricMistral Medium 3.5GPT-6.1 Sol
CatalogBenchmeasured by usChannel rules broken% of products 1.2%0.0%
CatalogBenchmeasured by usClaims to check% of products 25.6%0.0%
CatalogBenchmeasured by usContent quality% of checks 94.1%98.8%
CatalogBenchmeasured by usFailed outputs% of products 0.0%0.0%
CatalogBenchmeasured by usMissing UK information% of products 0.0%0.0%
CatalogBenchmeasured by usNot findable% of products 10.1%5.4%
CatalogBenchmeasured by usPublish-ready listings% of products 17.3%72.0%
CatalogBenchmeasured by usReliably publish-ready% of products 10.7%67.9%
CatalogBenchmeasured by usUnsupported claimsclaims per product 3.550.07
CatalogBenchmeasured by usUnsupported claims% of products 76.8%3.6%
CatalogBenchmeasured by usWrong attributes% of products 36.3%19.1%
CatalogBenchmeasured by usWrong category or variant% of products 13.1%1.8%
CatalogBenchmeasured by usChannel compliance% of products 100.0%100.0%
CatalogBenchmeasured by usChannel rules broken% of products 0.0%0.0%
CatalogBenchmeasured by usClaims to check% of products 3.6%0.0%
CatalogBenchmeasured by usConflicts caught% of conflicts 90.9%100.0%
CatalogBenchmeasured by usContent quality% of checks 94.1%98.6%
CatalogBenchmeasured by usCost per productUS dollars $0.0080$0.0084
CatalogBenchmeasured by usDecision accuracy% of decisions 92.6%98.7%
CatalogBenchmeasured by usFailed outputs% of products 0.0%0.0%
CatalogBenchmeasured by usField accuracy% of missing fields 91.7%94.4%
CatalogBenchmeasured by usInvented values% of filled values 0.0%2.2%
CatalogBenchmeasured by usMissing UK information% of products 0.0%0.0%
CatalogBenchmeasured by usNot findable% of products 14.3%4.8%
CatalogBenchmeasured by usPublish-ready listings% of products 29.8%78.0%
CatalogBenchmeasured by usReliably publish-ready% of products 17.9%75.0%
CatalogBenchmeasured by usUnsupported claims% of products 49.4%0.0%
CatalogBenchmeasured by usUnsupported claimsclaims per product 1.740.00
CatalogBenchmeasured by usWrong attributes% of products 30.9%16.7%
CatalogBenchmeasured by usWrong category or variant% of products 11.3%1.8%
DeckBenchmeasured by usAccurate decks% of tasks 0.0%100.0%
DeckBenchmeasured by usCaveat dropped% of tasks 0.0%0.0%
DeckBenchmeasured by usClean layout% of tasks 66.7%83.3%
DeckBenchmeasured by usCost per deckUS dollars $0.0448$0.15
DeckBenchmeasured by usDeck ratingrating 6491,412
DeckBenchmeasured by usDesign quality% of the maximum 43.6%71.4%
DeckBenchmeasured by usDraft figure quoted% of tasks 50.0%0.0%
DeckBenchmeasured by usFindings missing% of tasks 0.0%0.0%
DeckBenchmeasured by usHead-to-head win rate% of comparisons 18.6%74.6%
DeckBenchmeasured by usLayout defects% of tasks 33.3%16.7%
DeckBenchmeasured by usMisleading metric used% of tasks 0.0%0.0%
DeckBenchmeasured by usPresentable decks% of tasks 0.0%50.0%
DeckBenchmeasured by usRecommendation late or wrong% of tasks 50.0%0.0%
DeckBenchmeasured by usSlides needing work% of tasks 50.0%50.0%
DeckBenchmeasured by usUnsupported claims% of tasks 33.3%0.0%
DeckBenchmeasured by usUnsupported numbers% of tasks 0.0%0.0%
Artificial AnalysisreportedArtificial Analysis Intelligence Indexindex score 14.251.8
Artificial AnalysisreportedHumanity's Last Exam% of questions 13.8%52.9%
Artificial AnalysisreportedLong-context reasoning (AA-LCR)% of questions 69.3%84.0%
Artificial AnalysisreportedOutput speedtokens per second 16357.2
Artificial AnalysisreportedSciCode% of problems 40.2%55.8%
Artificial AnalysisreportedTerminal-Bench 4.0% of tasks 0.0%56.1%
Artificial AnalysisreportedTime to first answer tokenseconds 12.91.51

Bold green marks the better value on that metric. Where a source reports ranges that overlap, the difference may not be meaningful; see the benchmark page for ranges.

Quick answers

Is Mistral Medium 3.5 better than GPT-6.1 Sol?

It depends on the task. Among models available today, on our benchmarks and licensed sources, GPT-6.1 Sol ranks higher for product listings, decks from an analysis, agents and tool use, reasoning and knowledge and sticking to the facts.

Which is cheaper, Mistral Medium 3.5 or GPT-6.1 Sol?

Mistral Medium 3.5 costs $3.00 per million tokens (three input to one output), against $4.00 for GPT-6.1 Sol.

Which is faster, Mistral Medium 3.5 or GPT-6.1 Sol?

Mistral Medium 3.5 writes about 163 tokens a second on its usual API, against 57 for GPT-6.1 Sol (Artificial Analysis).

Run Mistral Medium 3.5 and GPT-6.1 Sol on your own prompt

Benchmarks aren't your data. Try both side by side in the playground, with the cost of every answer.