Compare / DeepSeek V4.1 Flash vs Mistral Large 4

DeepSeek V4.1 Flash vs Mistral Large 4

37 results from 2 sources that measured both models. Each row is in the source's own units; there is no overall winner.

DeepSeek V4.1 Flash
DeepSeek · profile
Mistral Large 4
Mistral · profile
Results compared
37
Shared sources
2
BenchmarkMetricDeepSeek V4.1 FlashMistral Large 4
CatalogBenchmeasured by usChannel rules broken% of products 1.8%1.2%
CatalogBenchmeasured by usClaims to check% of products 0.0%3.3%
CatalogBenchmeasured by usContent quality% of checks 96.2%95.9%
CatalogBenchmeasured by usFailed outputs% of products 0.6%0.0%
CatalogBenchmeasured by usMissing UK information% of products 0.0%0.0%
CatalogBenchmeasured by usNot findable% of products 3.6%3.0%
CatalogBenchmeasured by usPublish-ready listings% of products 47.0%10.7%
CatalogBenchmeasured by usReliably publish-ready% of products 23.2%1.8%
CatalogBenchmeasured by usUnsupported claims% of products 39.9%80.4%
CatalogBenchmeasured by usUnsupported claimsclaims per product 1.253.69
CatalogBenchmeasured by usWrong attributes% of products 20.2%17.3%
CatalogBenchmeasured by usWrong category or variant% of products 2.4%1.8%
CatalogBenchmeasured by usChannel compliance% of products 99.4%100.0%
CatalogBenchmeasured by usChannel rules broken% of products 3.0%3.0%
CatalogBenchmeasured by usClaims to check% of products 0.0%0.7%
CatalogBenchmeasured by usConflicts caught% of conflicts 100.0%100.0%
CatalogBenchmeasured by usContent quality% of checks 96.7%96.5%
CatalogBenchmeasured by usCost per productUS dollars $0.0061$0.0033
CatalogBenchmeasured by usDecision accuracy% of decisions 96.9%93.1%
CatalogBenchmeasured by usFailed outputs% of products 0.0%0.0%
CatalogBenchmeasured by usField accuracy% of missing fields 94.4%94.4%
CatalogBenchmeasured by usInvented values% of filled values 3.1%3.5%
CatalogBenchmeasured by usMissing UK information% of products 0.0%0.0%
CatalogBenchmeasured by usNot findable% of products 3.6%4.2%
CatalogBenchmeasured by usPublish-ready listings% of products 63.7%39.9%
CatalogBenchmeasured by usReliably publish-ready% of products 41.1%10.7%
CatalogBenchmeasured by usUnsupported claims% of products 13.1%33.3%
CatalogBenchmeasured by usUnsupported claimsclaims per product 0.540.91
CatalogBenchmeasured by usWrong attributes% of products 20.2%19.6%
CatalogBenchmeasured by usWrong category or variant% of products 3.0%1.2%
Artificial AnalysisreportedArtificial Analysis Intelligence Indexindex score 39.538.4
Artificial AnalysisreportedHumanity's Last Exam% of questions 39.2%35.0%
Artificial AnalysisreportedLong-context reasoning (AA-LCR)% of questions 84.0%81.3%
Artificial AnalysisreportedOutput speedtokens per second 213106
Artificial AnalysisreportedSciCode% of problems 51.9%54.2%
Artificial AnalysisreportedTerminal-Bench 4.0% of tasks 26.8%26.8%
Artificial AnalysisreportedTime to first answer tokenseconds 0.8919.8

Bold green marks the better value on that metric. Where a source reports ranges that overlap, the difference may not be meaningful; see the benchmark page for ranges.