BMC Bioinformatics (Feb 2019)

MultiDomainBenchmark: a multi-domain query and subject database suite

  • Hyrum D. Carroll,
  • John L. Spouge,
  • Mileidy Gonzalez

DOI
https://doi.org/10.1186/s12859-019-2660-5
Journal volume & issue
Vol. 20, no. 1
pp. 1 – 9

Abstract

Read online

Abstract Background Genetic sequence database retrieval benchmarks play an essential role in evaluating the performance of sequence searching tools. To date, all phylogenetically diverse benchmarks known to the authors include only query sequences with single protein domains. Domains are the primary building blocks of protein structure and function. Independently, each domain can fulfill a single function, but most proteins (>80% in Metazoa) exist as multi-domain proteins. Multiple domain units combine in various arrangements or architectures to create different functions and are often under evolutionary pressures to yield new ones. Thus, it is crucial to create gold standards reflecting the multi-domain complexity of real proteins to more accurately evaluate sequence searching tools. Description This work introduces MultiDomainBenchmark (MDB), a database suite of 412 curated multi-domain queries and 227,512 target sequences, representing at least 5108 species and 1123 phylogenetically divergent protein families, their relevancy annotation, and domain location. Here, we use the benchmark to evaluate the performance of two commonly used sequence searching tools, BLAST/PSI-BLAST and HMMER. Additionally, we introduce a novel classification technique for multi-domain proteins to evaluate how well an algorithm recovers a domain architecture. Conclusion MDB is publicly available at http://csc.columbusstate.edu/carroll/MDB/.

Keywords