STAR Protocols (Dec 2022)
Protocol to identify functional doppelgängers and verify biomedical gene expression data using doppelgangerIdentifier
Abstract
Summary: Functional doppelgängers (FDs) are independently derived sample pairs that confound machine learning model (ML) performance when assorted across training and validation sets. Here, we detail the use of doppelgangerIdentifier (DI), providing software installation, data preparation, doppelgänger identification, and functional testing steps. We demonstrate examples with biomedical gene expression data. We also provide guidelines for the selection of user-defined function arguments.For complete details on the use and execution of this protocol, please refer to Wang et al. (2022). : Publisher’s note: Undertaking any experimental protocol requires adherence to local institutional guidelines for laboratory safety and ethics.