RepAHR: an improved approach for de novo repeat identification by assembly of the high-frequency reads

Xingyu Liao; Xin Gao; Xiankai Zhang; Fang-Xiang Wu; Jianxin Wang

doi:10.1186/s12859-020-03779-w

BMC Bioinformatics (Oct 2020)

RepAHR: an improved approach for de novo repeat identification by assembly of the high-frequency reads

Xingyu Liao,
Xin Gao,
Xiankai Zhang,
Fang-Xiang Wu,
Jianxin Wang

Affiliations

Xingyu Liao: School of Computer Science and Engineering, Central South University
Xin Gao: Computational Bioscience Research Center, Computer, Electrical and Mathematical Sciences and Engineering Division, King Abdullah University of Science and Technology (KAUST)
Xiankai Zhang: School of Computer Science and Engineering, Central South University
Fang-Xiang Wu: Biomedical Engineering and Department of Mechanical Engineering, University of Saskatchewan
Jianxin Wang: School of Computer Science and Engineering, Central South University

DOI: https://doi.org/10.1186/s12859-020-03779-w
Journal volume & issue: Vol. 21, no. 1
pp. 1 – 24

Abstract

Read online

Abstract Background Repetitive sequences account for a large proportion of eukaryotes genomes. Identification of repetitive sequences plays a significant role in many applications, such as structural variation detection and genome assembly. Many existing de novo repeat identification pipelines or tools make use of assembly of the high-frequency k-mers to obtain repeats. However, a certain degree of sequence coverage is required for assemblers to get the desired assemblies. On the other hand, assemblers cut the reads into shorter k-mers for assembly, which may destroy the structure of the repetitive regions. For the above reasons, it is difficult to obtain complete and accurate repetitive regions in the genome by using existing tools. Results In this study, we present a new method called RepAHR for de novo repeat identification by assembly of the high-frequency reads. Firstly, RepAHR scans next-generation sequencing (NGS) reads to find the high-frequency k-mers. Secondly, RepAHR filters the high-frequency reads from whole NGS reads according to certain rules based on the high-frequency k-mer. Finally, the high-frequency reads are assembled to generate repeats by using SPAdes, which is considered as an outstanding genome assembler with NGS sequences. Conlusions We test RepAHR on five data sets, and the experimental results show that RepAHR outperforms RepARK and REPdenovo for detecting repeats in terms of N50, reference alignment ratio, coverage ratio of reference, mask ratio of Repbase and some other metrics.

Published in BMC Bioinformatics

ISSN: 1471-2105 (Online)
Publisher: BMC
Country of publisher: United Kingdom
LCC subjects: Medicine: Medicine (General): Computer applications to medicine. Medical informatics; Science: Biology (General)
Website: http://www.biomedcentral.com/bmcbioinformatics/

About the journal

Abstract

Keywords