Information (Oct 2016)

A Survey on Data Compression Methods for Biological Sequences

  • Morteza Hosseini,
  • Diogo Pratas,
  • Armando J. Pinho

DOI
https://doi.org/10.3390/info7040056
Journal volume & issue
Vol. 7, no. 4
p. 56

Abstract

Read online

The ever increasing growth of the production of high-throughput sequencing data poses a serious challenge to the storage, processing and transmission of these data. As frequently stated, it is a data deluge. Compression is essential to address this challenge—it reduces storage space and processing costs, along with speeding up data transmission. In this paper, we provide a comprehensive survey of existing compression approaches, that are specialized for biological data, including protein and DNA sequences. Also, we devote an important part of the paper to the approaches proposed for the compression of different file formats, such as FASTA, as well as FASTQ and SAM/BAM, which contain quality scores and metadata, in addition to the biological sequences. Then, we present a comparison of the performance of several methods, in terms of compression ratio, memory usage and compression/decompression time. Finally, we present some suggestions for future research on biological data compression.

Keywords