Compare / DeepSeek V4 Pro 0423 vs Mistral Large 4

DeepSeek V4 Pro 0423 vs Mistral Large 4

32 results from 3 sources that measured both models. Each row is in the source's own units; there is no overall winner.

DeepSeek V4 Pro 0423
DeepSeek · profile
Mistral Large 4
Mistral · profile
Results compared
32
Shared sources
3
BenchmarkMetricDeepSeek V4 Pro 0423Mistral Large 4
DeckBenchmeasured by usAccurate decks% of tasks 0.0%0.0%
DeckBenchmeasured by usCaveat dropped% of tasks 16.7%16.7%
DeckBenchmeasured by usClean layout% of tasks 33.3%33.3%
DeckBenchmeasured by usCost per deckUS dollars $0.0465$0.0148
DeckBenchmeasured by usDeck ratingrating 615515
DeckBenchmeasured by usDesign quality% of the maximum 40.9%30.5%
DeckBenchmeasured by usDraft figure quoted% of tasks 33.3%16.7%
DeckBenchmeasured by usFindings missing% of tasks 16.7%0.0%
DeckBenchmeasured by usHead-to-head win rate% of comparisons 18.2%10.7%
DeckBenchmeasured by usLayout defects% of tasks 66.7%66.7%
DeckBenchmeasured by usMisleading metric used% of tasks 0.0%0.0%
DeckBenchmeasured by usPresentable decks% of tasks 0.0%0.0%
DeckBenchmeasured by usRecommendation late or wrong% of tasks 66.7%83.3%
DeckBenchmeasured by usSlides needing work% of tasks 83.3%100.0%
DeckBenchmeasured by usUnsupported claims% of tasks 66.7%50.0%
DeckBenchmeasured by usUnsupported numbers% of tasks 0.0%33.3%
ROASBenchmeasured by usAudience scorescore out of 100 44.855.8
ROASBenchmeasured by usBusiness scorescore out of 100 23.834.2
ROASBenchmeasured by usConsistency scorescore out of 100 9.116.7
ROASBenchmeasured by usContribution profitsimulated US dollars $301,289$533,401
ROASBenchmeasured by usCost of a runUS dollars $0.20$0.0657
ROASBenchmeasured by usMonths over budgetmonths of 12 00
ROASBenchmeasured by usOverall scorescore out of 100 28.237.0
ROASBenchmeasured by usPlanning scorescore out of 100 54.055.0
ROASBenchmeasured by usReturn on ad spendprofit per $1 spent $0.26$0.45
Artificial AnalysisreportedArtificial Analysis Intelligence Indexindex score 36.038.4
Artificial AnalysisreportedHumanity's Last Exam% of questions 41.0%35.0%
Artificial AnalysisreportedLong-context reasoning (AA-LCR)% of questions 80.3%81.3%
Artificial AnalysisreportedOutput speedtokens per second 117106
Artificial AnalysisreportedSciCode% of problems 51.0%54.2%
Artificial AnalysisreportedTerminal-Bench 4.0% of tasks 14.1%26.8%
Artificial AnalysisreportedTime to first answer tokenseconds 0.9019.8

Bold green marks the better value on that metric. Where a source reports ranges that overlap, the difference may not be meaningful; see the benchmark page for ranges.