Accurate prediction of protein structural class.

Xia-Yu Xia; Meng Ge; Zhi-Xin Wang; Xian-Ming Pan

doi:10.1371/journal.pone.0037653

PLoS ONE (Jan 2012)

Accurate prediction of protein structural class.

Xia-Yu Xia,
Meng Ge,
Zhi-Xin Wang,
Xian-Ming Pan

Affiliations

Xia-Yu Xia
Meng Ge
Zhi-Xin Wang
Xian-Ming Pan

DOI: https://doi.org/10.1371/journal.pone.0037653
Journal volume & issue: Vol. 7, no. 6
p. e37653

Abstract

Read online

Because of the increasing gap between the data from sequencing and structural genomics, the accurate prediction of the structural class of a protein domain solely from the primary sequence has remained a challenging problem in structural biology. Traditional sequence-based predictors generally select several sequence features and then feed them directly into a classification program to identify the structural class. The current best sequence-based predictor achieved an overall accuracy of 74.1% when tested on a widely used, non-homologous benchmark dataset 25PDB. In the present work, we built a multiple linear regression (MLR) model to convert the 440-dimensional (440D) sequence feature vector extracted from the Position Specific Scoring Matrix (PSSM) of a protein domain to a 4-dimensinal (4D) structural feature vector, which could then be used to predict the four major structural classes. We performed 10-fold cross-validation and jackknife tests of the method on a large non-homologous dataset containing 8,244 domains distributed among the four major classes. The performance of our approach outperformed all of the existing sequence-based methods and had an overall accuracy of 83.1%, which is even higher than the results of those predicted secondary structure-based methods.

Published in PLoS ONE

ISSN: 1932-6203 (Online)
Publisher: Public Library of Science (PLoS)
Country of publisher: United States
LCC subjects: Medicine; Science
Website: https://journals.plos.org/plosone/

About the journal