A pipeline for sample tagging of whole genome bisulfite sequencing data using genotypes of whole genome sequencing

Zhe Xu; Si Cheng; Xin Qiu; Xiaoqi Wang; Qiuwen Hu; Yanfeng Shi; Yang Liu; Jinxi Lin; Jichao Tian; Yongfei Peng; Yong Jiang; Yadong Yang; Jianwei Ye; Yilong Wang; Xia Meng; Zixiao Li; Hao Li; Yongjun Wang

doi:10.1186/s12864-023-09413-2

BMC Genomics (Jun 2023)

A pipeline for sample tagging of whole genome bisulfite sequencing data using genotypes of whole genome sequencing

Zhe Xu,
Si Cheng,
Xin Qiu,
Xiaoqi Wang,
Qiuwen Hu,
Yanfeng Shi,
Yang Liu,
Jinxi Lin,
Jichao Tian,
Yongfei Peng,
Yong Jiang,
Yadong Yang,
Jianwei Ye,
Yilong Wang,
Xia Meng,
Zixiao Li,
Hao Li,
Yongjun Wang

Affiliations

Zhe Xu: Department of Neurology, Beijing Tiantan Hospital, Capital Medical University
Si Cheng: Department of Neurology, Beijing Tiantan Hospital, Capital Medical University
Xin Qiu: Department of Neurology, Beijing Tiantan Hospital, Capital Medical University
Xiaoqi Wang: BioChain (Beijing) Science and Technology, Inc, Economic and Technological Development Area
Qiuwen Hu: BioChain (Beijing) Science and Technology, Inc, Economic and Technological Development Area
Yanfeng Shi: Department of Neurology, Beijing Tiantan Hospital, Capital Medical University
Yang Liu: Department of Neurology, Beijing Tiantan Hospital, Capital Medical University
Jinxi Lin: Department of Neurology, Beijing Tiantan Hospital, Capital Medical University
Jichao Tian: BioChain (Beijing) Science and Technology, Inc, Economic and Technological Development Area
Yongfei Peng: BioChain (Beijing) Science and Technology, Inc, Economic and Technological Development Area
Yong Jiang: Department of Neurology, Beijing Tiantan Hospital, Capital Medical University
Yadong Yang: BioChain (Beijing) Science and Technology, Inc, Economic and Technological Development Area
Jianwei Ye: BioChain (Beijing) Science and Technology, Inc, Economic and Technological Development Area
Yilong Wang: Department of Neurology, Beijing Tiantan Hospital, Capital Medical University
Xia Meng: Department of Neurology, Beijing Tiantan Hospital, Capital Medical University
Zixiao Li: Department of Neurology, Beijing Tiantan Hospital, Capital Medical University
Hao Li: Department of Neurology, Beijing Tiantan Hospital, Capital Medical University
Yongjun Wang: Department of Neurology, Beijing Tiantan Hospital, Capital Medical University

DOI: https://doi.org/10.1186/s12864-023-09413-2
Journal volume & issue: Vol. 24, no. 1
pp. 1 – 13

Abstract

Read online

Abstract Background In large-scale high-throughput sequencing projects and biobank construction, sample tagging is essential to prevent sample mix-ups. Despite the availability of fingerprint panels for DNA data, little research has been conducted on sample tagging of whole genome bisulfite sequencing (WGBS) data. This study aims to construct a pipeline and identify applicable fingerprint panels to address this problem. Results Using autosome-wide A/T polymorphic single nucleotide variants (SNVs) obtained from whole genome sequencing (WGS) and WGBS of individuals from the Third China National Stroke Registry, we designed a fingerprint panel and constructed an optimized pipeline for tagging WGBS data. This pipeline used Bis-SNP to call genotypes from the WGBS data, and optimized genotype comparison by eliminating wildtype homozygous and missing genotypes, and retaining variants with identical genomic coordinates and reference/alternative alleles. WGS-based and WGBS-based genotypes called from identical or different samples were extensively compared using hap.py. In the first batch of 94 samples, the genotype consistency rates were between 71.01%-84.23% and 51.43%-60.50% for the matched and mismatched WGS and WGBS data using the autosome-wide A/T polymorphic SNV panel. This capability to tag WGBS data was validated among the second batch of 240 samples, with genotype consistency rates ranging from 70.61%-84.65% to 49.58%-61.42% for the matched and mismatched data, respectively. We also determined that the number of genetic variants required to correctly tag WGBS data was on the order of thousands through testing six fingerprint panels with different orders for the number of variants. Additionally, we affirmed this result with two self-designed panels of 1351 and 1278 SNVs, respectively. Furthermore, this study confirmed that using the number of genetic variants with identical coordinates and ref/alt alleles, or identical genotypes could not correctly tag WGBS data. Conclusion This study proposed an optimized pipeline, applicable fingerprint panels, and a lower boundary for the number of fingerprint genetic variants needed for correct sample tagging of WGBS data, which are valuable for tagging WGBS data and integrating multi-omics data for biobanks.

Published in BMC Genomics

ISSN: 1471-2164 (Online)
Publisher: BMC
Country of publisher: United Kingdom
LCC subjects: Technology: Chemical technology: Biotechnology; Science: Biology (General): Genetics
Website: http://bmcgenomics.biomedcentral.com

About the journal

Abstract

Keywords