PLoS ONE (Mar 2011)

A practical comparison of de novo genome assembly software tools for next-generation sequencing technologies.

  • Wenyu Zhang,
  • Jiajia Chen,
  • Yang Yang,
  • Yifei Tang,
  • Jing Shang,
  • Bairong Shen

DOI
https://doi.org/10.1371/journal.pone.0017915
Journal volume & issue
Vol. 6, no. 3
p. e17915

Abstract

Read online

The advent of next-generation sequencing technologies is accompanied with the development of many whole-genome sequence assembly methods and software, especially for de novo fragment assembly. Due to the poor knowledge about the applicability and performance of these software tools, choosing a befitting assembler becomes a tough task. Here, we provide the information of adaptivity for each program, then above all, compare the performance of eight distinct tools against eight groups of simulated datasets from Solexa sequencing platform. Considering the computational time, maximum random access memory (RAM) occupancy, assembly accuracy and integrity, our study indicate that string-based assemblers, overlap-layout-consensus (OLC) assemblers are well-suited for very short reads and longer reads of small genomes respectively. For large datasets of more than hundred millions of short reads, De Bruijn graph-based assemblers would be more appropriate. In terms of software implementation, string-based assemblers are superior to graph-based ones, of which SOAPdenovo is complex for the creation of configuration file. Our comparison study will assist researchers in selecting a well-suited assembler and offer essential information for the improvement of existing assemblers or the developing of novel assemblers.