Compare / Mistral Large 4 vs Grok 4.7

Mistral Large 4 vs Grok 4.7

Which is cheaper, which is more capable, and which does better at each kind of work, from 62 results on 4 sources that measured both models.

Mistral Large 4
Mistral · profile
Grok 4.7
xAI · profile
Results compared
62
Shared sources
4

Mistral Large 4 or Grok 4.7? The short answer

  • Mistral Large 4 is cheaper: $1.03 against $3.00 per million tokens, 2.9 times less.
  • Grok 4.7 scores higher on the Artificial Analysis Intelligence Index: 46.4 against 38.4.
  • Mistral Large 4 writes faster: 106 tokens a second against 68.
  • On our business benchmarks, Mistral Large 4 ranks higher for marketing planning, agents and tool use and speed under pressure.
  • On our business benchmarks, Grok 4.7 ranks higher for product listings, decks from an analysis and reasoning and knowledge.

At a glance

Mistral Large 4Grok 4.7
Price per million tokens$0.68 in · $2.09 out$2.00 in · $6.00 out
Blended (3 in : 1 out)$1.03$3.00
Cheapest host, blended–$3.00 (xAI)
Intelligence Index38.446.4
Spring Prompt overall3143
Output speed, tokens a second10668
Context window524,288 tokens500,000 tokens
WeightsClosedClosed
Developer based inFrancethe US

Prices are the developer's list price where we have read it, otherwise the typical host on OpenRouter; see every model's price. Intelligence and speed from Artificial Analysis.

Which is better for what

Each model's rank among models available today, for each kind of work, from our own benchmarks and licensed sources. The better rank is in green.

Use caseMistral Large 4Grok 4.7
Product listingsTurning a sparse product feed and photos into listings that can go live 17 of 19 4 of 19
Decks from an analysisTurning a finished analysis into a deck you could present as it is 18 of 18 9 of 18
Marketing planningPlanning a year of ad spend without overspending 10 of 18 11 of 18
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls 20 of 161 26 of 161
Reasoning and knowledgeHard questions across science, maths and general knowledge 43 of 195 16 of 195
WritingWhat people prefer in blind comparisons, and judged writing quality not measured 63 of 187
Sticking to the factsSummarising without inventing things, and factual answers not measured 112 of 153
Speed under pressureGood decisions against a real clock (fast chess) 13 of 21 22 of 22

Their slides, side by side

The same brief for both: Manchester store investment review for Hearthside Coffee, an invented company. First slide of each deck from DeckBench.

First slide of Mistral Large 4's deck for Hearthside Coffee
Mistral Large 4 · deck rating 515 · needs work
First slide of Grok 4.7's deck for Hearthside Coffee
Grok 4.7 · deck rating 1,242 · presentable as it is

Every slide, side by side →

From our research

Every shared result

BenchmarkMetricMistral Large 4Grok 4.7
CatalogBenchmeasured by usChannel rules broken% of products 1.2%0.0%
CatalogBenchmeasured by usClaims to check% of products 3.3%0.0%
CatalogBenchmeasured by usContent quality% of checks 95.9%98.1%
CatalogBenchmeasured by usFailed outputs% of products 0.0%0.0%
CatalogBenchmeasured by usMissing UK information% of products 0.0%0.0%
CatalogBenchmeasured by usNot findable% of products 3.0%4.2%
CatalogBenchmeasured by usPublish-ready listings% of products 10.7%58.9%
CatalogBenchmeasured by usReliably publish-ready% of products 1.8%33.9%
CatalogBenchmeasured by usUnsupported claimsclaims per product 3.690.71
CatalogBenchmeasured by usUnsupported claims% of products 80.4%26.8%
CatalogBenchmeasured by usWrong attributes% of products 17.3%15.5%
CatalogBenchmeasured by usWrong category or variant% of products 1.8%2.4%
CatalogBenchmeasured by usChannel compliance% of products 100.0%100.0%
CatalogBenchmeasured by usChannel rules broken% of products 3.0%0.0%
CatalogBenchmeasured by usClaims to check% of products 0.7%0.0%
CatalogBenchmeasured by usConflicts caught% of conflicts 100.0%100.0%
CatalogBenchmeasured by usContent quality% of checks 96.5%97.6%
CatalogBenchmeasured by usCost per productUS dollars $0.0033$0.0237
CatalogBenchmeasured by usDecision accuracy% of decisions 93.1%97.3%
CatalogBenchmeasured by usFailed outputs% of products 0.0%0.0%
CatalogBenchmeasured by usField accuracy% of missing fields 94.4%92.9%
CatalogBenchmeasured by usInvented values% of filled values 3.5%0.5%
CatalogBenchmeasured by usMissing UK information% of products 0.0%0.0%
CatalogBenchmeasured by usNot findable% of products 4.2%5.4%
CatalogBenchmeasured by usPublish-ready listings% of products 39.9%70.8%
CatalogBenchmeasured by usReliably publish-ready% of products 10.7%57.1%
CatalogBenchmeasured by usUnsupported claims% of products 33.3%11.3%
CatalogBenchmeasured by usUnsupported claimsclaims per product 0.910.49
CatalogBenchmeasured by usWrong attributes% of products 19.6%15.5%
CatalogBenchmeasured by usWrong category or variant% of products 1.2%2.4%
DeckBenchmeasured by usAccurate decks% of tasks 0.0%66.7%
DeckBenchmeasured by usCaveat dropped% of tasks 16.7%16.7%
DeckBenchmeasured by usClean layout% of tasks 33.3%83.3%
DeckBenchmeasured by usCost per deckUS dollars $0.0148$0.25
DeckBenchmeasured by usDeck ratingrating 5151,242
DeckBenchmeasured by usDesign quality% of the maximum 30.5%76.4%
DeckBenchmeasured by usDraft figure quoted% of tasks 16.7%16.7%
DeckBenchmeasured by usFindings missing% of tasks 0.0%0.0%
DeckBenchmeasured by usHead-to-head win rate% of comparisons 10.7%68.7%
DeckBenchmeasured by usLayout defects% of tasks 66.7%16.7%
DeckBenchmeasured by usMisleading metric used% of tasks 0.0%0.0%
DeckBenchmeasured by usPresentable decks% of tasks 0.0%33.3%
DeckBenchmeasured by usRecommendation late or wrong% of tasks 83.3%16.7%
DeckBenchmeasured by usSlides needing work% of tasks 100.0%16.7%
DeckBenchmeasured by usUnsupported claims% of tasks 50.0%0.0%
DeckBenchmeasured by usUnsupported numbers% of tasks 33.3%0.0%
ROASBenchmeasured by usAudience scorescore out of 100 55.858.2
ROASBenchmeasured by usBusiness scorescore out of 100 34.226.9
ROASBenchmeasured by usConsistency scorescore out of 100 16.730.5
ROASBenchmeasured by usContribution profitsimulated US dollars $533,401$187,636
ROASBenchmeasured by usCost of a runUS dollars $0.0657$0.42
ROASBenchmeasured by usMonths over budgetmonths of 12 00
ROASBenchmeasured by usOverall scorescore out of 100 37.036.6
ROASBenchmeasured by usPlanning scorescore out of 100 55.054.6
ROASBenchmeasured by usReturn on ad spendprofit per $1 spent $0.45$0.23
Artificial AnalysisreportedArtificial Analysis Intelligence Indexindex score 38.446.4
Artificial AnalysisreportedHumanity's Last Exam% of questions 35.0%43.1%
Artificial AnalysisreportedLong-context reasoning (AA-LCR)% of questions 81.3%80.0%
Artificial AnalysisreportedOutput speedtokens per second 10668.4
Artificial AnalysisreportedSciCode% of problems 54.2%57.8%
Artificial AnalysisreportedTerminal-Bench 4.0% of tasks 26.8%25.8%
Artificial AnalysisreportedTime to first answer tokenseconds 19.84.25

Bold green marks the better value on that metric. Where a source reports ranges that overlap, the difference may not be meaningful; see the benchmark page for ranges.

Quick answers

Is Mistral Large 4 better than Grok 4.7?

It depends on the task. Among models available today, on our benchmarks and licensed sources, Mistral Large 4 ranks higher for marketing planning, agents and tool use and speed under pressure; Grok 4.7 ranks higher for product listings, decks from an analysis and reasoning and knowledge.

Which is cheaper, Mistral Large 4 or Grok 4.7?

Mistral Large 4 costs $1.03 per million tokens (three input to one output), against $3.00 for Grok 4.7.

Which is faster, Mistral Large 4 or Grok 4.7?

Mistral Large 4 writes about 106 tokens a second on its usual API, against 68 for Grok 4.7 (Artificial Analysis).

Run Mistral Large 4 and Grok 4.7 on your own prompt

Benchmarks aren't your data. Try both side by side in the playground, with the cost of every answer.