Open-Source Sequence Clustering Methods Improve the State Of the Art

Evguenia Kopylova; Jose A. Navas-Molina; Céline Mercier; Zhenjiang Zech Xu; Frédéric Mahé; Yan He; Hong-Wei Zhou; Torbjørn Rognes; J. Gregory Caporaso; Rob Knight

doi:10.1128/mSystems.00003-15

mSystems (Feb 2016)

Open-Source Sequence Clustering Methods Improve the State Of the Art

Evguenia Kopylova,
Jose A. Navas-Molina,
Céline Mercier,
Zhenjiang Zech Xu,
Frédéric Mahé,
Yan He,
Hong-Wei Zhou,
Torbjørn Rognes,
J. Gregory Caporaso,
Rob Knight

Affiliations

Evguenia Kopylova: Department of Pediatrics, UCSD School of Medicine, La Jolla, California, USA
Jose A. Navas-Molina: Department of Pediatrics, UCSD School of Medicine, La Jolla, California, USA
Céline Mercier: Laboratoire d'Ecologie Alpine (LECA), CNRS UMR 5553, Université Grenoble Alpes, Grenoble, France
Zhenjiang Zech Xu: Department of Pediatrics, UCSD School of Medicine, La Jolla, California, USA
Frédéric Mahé: Department of Ecology, University of Kaiserslautern, Kaiserslautern, Germany
Yan He: Department of Environmental Health, State Key Laboratory of Organ Failure Research, Guangdong Provincial Key Laboratory of Tropical Disease Research, School of Public Health and Tropical Medicine, Southern Medical University, Guangzhou, Guangdong, China
Hong-Wei Zhou: Department of Environmental Health, State Key Laboratory of Organ Failure Research, Guangdong Provincial Key Laboratory of Tropical Disease Research, School of Public Health and Tropical Medicine, Southern Medical University, Guangzhou, Guangdong, China
Torbjørn Rognes: Department of Informatics, University of Oslo, Oslo, Norway
J. Gregory Caporaso: Department of Biological Sciences, Northern Arizona University, Flagstaff, Arizona, USA
Rob Knight: Department of Pediatrics, UCSD School of Medicine, La Jolla, California, USA

DOI: https://doi.org/10.1128/mSystems.00003-15
Journal volume & issue: Vol. 1, no. 1

Abstract

Read online

ABSTRACT Sequence clustering is a common early step in amplicon-based microbial community analysis, when raw sequencing reads are clustered into operational taxonomic units (OTUs) to reduce the run time of subsequent analysis steps. Here, we evaluated the performance of recently released state-of-the-art open-source clustering software products, namely, OTUCLUST, Swarm, SUMACLUST, and SortMeRNA, against current principal options (UCLUST and USEARCH) in QIIME, hierarchical clustering methods in mothur, and USEARCH’s most recent clustering algorithm, UPARSE. All the latest open-source tools showed promising results, reporting up to 60% fewer spurious OTUs than UCLUST, indicating that the underlying clustering algorithm can vastly reduce the number of these derived OTUs. Furthermore, we observed that stringent quality filtering, such as is done in UPARSE, can cause a significant underestimation of species abundance and diversity, leading to incorrect biological results. Swarm, SUMACLUST, and SortMeRNA have been included in the QIIME 1.9.0 release. IMPORTANCE Massive collections of next-generation sequencing data call for fast, accurate, and easily accessible bioinformatics algorithms to perform sequence clustering. A comprehensive benchmark is presented, including open-source tools and the popular USEARCH suite. Simulated, mock, and environmental communities were used to analyze sensitivity, selectivity, species diversity (alpha and beta), and taxonomic composition. The results demonstrate that recent clustering algorithms can significantly improve accuracy and preserve estimated diversity without the application of aggressive filtering. Moreover, these tools are all open source, apply multiple levels of multithreading, and scale to the demands of modern next-generation sequencing data, which is essential for the analysis of massive multidisciplinary studies such as the Earth Microbiome Project (EMP) (J. A. Gilbert, J. K. Jansson, and R. Knight, BMC Biol 12:69, 2014, http://dx.doi.org/10.1186/s12915-014-0069-1 ).

Published in mSystems

ISSN: 2379-5077 (Online)
Publisher: American Society for Microbiology
Country of publisher: United States
LCC subjects: Science: Microbiology
Website: https://journals.asm.org/journal/msystems

About the journal

Abstract

Keywords