Using penalized regression to predict phenotype from SNP data

Svetlana Cherlin; Richard A. J. Howey; Heather J. Cordell

doi:10.1186/s12919-018-0149-2

BMC Proceedings (Sep 2018)

Using penalized regression to predict phenotype from SNP data

Svetlana Cherlin,
Richard A. J. Howey,
Heather J. Cordell

Affiliations

Svetlana Cherlin: Institute of Genetic Medicine, Newcastle University, International Centre for Life
Richard A. J. Howey: Institute of Genetic Medicine, Newcastle University, International Centre for Life
Heather J. Cordell: Institute of Genetic Medicine, Newcastle University, International Centre for Life

DOI: https://doi.org/10.1186/s12919-018-0149-2
Journal volume & issue: Vol. 12, no. S9
pp. 223 – 228

Abstract

Read online

Abstract Background In a typical genome-enabled prediction problem there are many more predictor variables than response variables. This prohibits the application of multiple linear regression, because the unique ordinary least squares estimators of the regression coefficients are not defined. To overcome this problem, penalized regression methods have been proposed, aiming at shrinking the coefficients toward zero. Methods We explore prediction of phenotype from single nucleotide polymorphism (SNP) data in the GAW20 data set using a penalized regression approach (LASSO [least absolute shrinkage and selection operator] regression). We use 10-fold cross-validation to assess predictive performance and 10-fold nested cross-validation to specify a penalty parameter. Results By analyzing approximately 600,000 SNPs we find that, when the sample size comprises a few hundred individuals, SNP effects are heavily penalized, resulting in a poor predictive performance. Increasing the sample size to a few thousand individuals results in a much smaller penalization of the true effects, thus greatly improving the prediction. Conclusions LASSO regression results in a heavy shrinkage of the regression coefficients, and also requires large sample sizes (several thousand individuals) to achieve good prediction.

Published in BMC Proceedings

ISSN: 1753-6561 (Online)
Publisher: BMC
Country of publisher: United Kingdom
LCC subjects: Medicine; Science
Website: http://www.biomedcentral.com/bmcproc/

About the journal