Systematic evaluation with practical guidelines for single-cell and spatially resolved transcriptomics data simulation under multiple scenarios

Hongrui Duo; Yinghong Li; Yang Lan; Jingxin Tao; Qingxia Yang; Yingxue Xiao; Jing Sun; Lei Li; Xiner Nie; Xiaoxi Zhang; Guizhao Liang; Mingwei Liu; Youjin Hao; Bo Li

doi:10.1186/s13059-024-03290-y

Genome Biology (Jun 2024)

Systematic evaluation with practical guidelines for single-cell and spatially resolved transcriptomics data simulation under multiple scenarios

Hongrui Duo,
Yinghong Li,
Yang Lan,
Jingxin Tao,
Qingxia Yang,
Yingxue Xiao,
Jing Sun,
Lei Li,
Xiner Nie,
Xiaoxi Zhang,
Guizhao Liang,
Mingwei Liu,
Youjin Hao,
Bo Li

Affiliations

Hongrui Duo: College of Life Sciences, Chongqing Normal University
Yinghong Li: Chongqing Key Laboratory of Big Data for Bio Intelligence, Chongqing University of Posts and Telecommunications
Yang Lan: Institute of Pathology and Southwest Cancer Center, Southwest Hospital, Army Medical University
Jingxin Tao: College of Life Sciences, Chongqing Normal University
Qingxia Yang: Zhejiang Provincial Key Laboratory of Precision Diagnosis and Therapy for Major Gynecological Diseases, Women’s Hospital, Zhejiang University School of Medicine
Yingxue Xiao: College of Life Sciences, Chongqing Normal University
Jing Sun: College of Life Sciences, Chongqing Normal University
Lei Li: College of Life Sciences, Chongqing Normal University
Xiner Nie: Key Laboratory of Biorheological Science and Technology, Ministry of Education, Bioengineering College, Chongqing University
Xiaoxi Zhang: College of Life Sciences, Chongqing Normal University
Guizhao Liang: Key Laboratory of Biorheological Science and Technology, Ministry of Education, Bioengineering College, Chongqing University
Mingwei Liu: Key Laboratory of Clinical Laboratory Diagnostics, College of Laboratory Medicine, Chongqing Medical University
Youjin Hao: College of Life Sciences, Chongqing Normal University
Bo Li: College of Life Sciences, Chongqing Normal University

DOI: https://doi.org/10.1186/s13059-024-03290-y
Journal volume & issue: Vol. 25, no. 1
pp. 1 – 29

Abstract

Read online

Abstract Background Single-cell RNA sequencing (scRNA-seq) and spatially resolved transcriptomics (SRT) have led to groundbreaking advancements in life sciences. To develop bioinformatics tools for scRNA-seq and SRT data and perform unbiased benchmarks, data simulation has been widely adopted by providing explicit ground truth and generating customized datasets. However, the performance of simulation methods under multiple scenarios has not been comprehensively assessed, making it challenging to choose suitable methods without practical guidelines. Results We systematically evaluated 49 simulation methods developed for scRNA-seq and/or SRT data in terms of accuracy, functionality, scalability, and usability using 152 reference datasets derived from 24 platforms. SRTsim, scDesign3, ZINB-WaVE, and scDesign2 have the best accuracy performance across various platforms. Unexpectedly, some methods tailored to scRNA-seq data have potential compatibility for simulating SRT data. Lun, SPARSim, and scDesign3-tree outperform other methods under corresponding simulation scenarios. Phenopath, Lun, Simple, and MFA yield high scalability scores but they cannot generate realistic simulated data. Users should consider the trade-offs between method accuracy and scalability (or functionality) when making decisions. Additionally, execution errors are mainly caused by failed parameter estimations and appearance of missing or infinite values in calculations. We provide practical guidelines for method selection, a standard pipeline Simpipe ( https://github.com/duohongrui/simpipe ; https://doi.org/10.5281/zenodo.11178409 ), and an online tool Simsite ( https://www.ciblab.net/software/simshiny/ ) for data simulation. Conclusions No method performs best on all criteria, thus a good-yet-not-the-best method is recommended if it solves problems effectively and reasonably. Our comprehensive work provides crucial insights for developers on modeling gene expression data and fosters the simulation process for users.

Published in Genome Biology

ISSN: 1474-760X (Online)
Publisher: BMC
Country of publisher: United Kingdom
LCC subjects: Science: Biology (General): Genetics
Website: https://genomebiology.biomedcentral.com/

About the journal

Abstract

Keywords