Model-based clustering of DNA methylation array data: a recursive-partitioning algorithm for high-dimensional data arising as a mixture of beta distributions

Wiemels Joseph; Nelson Heather H; Wrensch Margaret; Karagas Margaret R; Marsit Carmen J; Yeh Ru-Fang; Christensen Brock C; Houseman E Andres; Zheng Shichun; Wiencke John K; Kelsey Karl T

doi:10.1186/1471-2105-9-365

BMC Bioinformatics (Sep 2008)

Model-based clustering of DNA methylation array data: a recursive-partitioning algorithm for high-dimensional data arising as a mixture of beta distributions

Wiemels Joseph,
Nelson Heather H,
Wrensch Margaret,
Karagas Margaret R,
Marsit Carmen J,
Yeh Ru-Fang,
Christensen Brock C,
Houseman E Andres,
Zheng Shichun,
Wiencke John K,
Kelsey Karl T

Affiliations

Wiemels Joseph
Nelson Heather H
Wrensch Margaret
Karagas Margaret R
Marsit Carmen J
Yeh Ru-Fang
Christensen Brock C
Houseman E Andres
Zheng Shichun
Wiencke John K
Kelsey Karl T

DOI: https://doi.org/10.1186/1471-2105-9-365
Journal volume & issue: Vol. 9, no. 1
p. 365

Abstract

Read online

Abstract Background Epigenetics is the study of heritable changes in gene function that cannot be explained by changes in DNA sequence. One of the most commonly studied epigenetic alterations is cytosine methylation, which is a well recognized mechanism of epigenetic gene silencing and often occurs at tumor suppressor gene loci in human cancer. Arrays are now being used to study DNA methylation at a large number of loci; for example, the Illumina GoldenGate platform assesses DNA methylation at 1505 loci associated with over 800 cancer-related genes. Model-based cluster analysis is often used to identify DNA methylation subgroups in data, but it is unclear how to cluster DNA methylation data from arrays in a scalable and reliable manner. Results We propose a novel model-based recursive-partitioning algorithm to navigate clusters in a beta mixture model. We present simulations that show that the method is more reliable than competing nonparametric clustering approaches, and is at least as reliable as conventional mixture model methods. We also show that our proposed method is more computationally efficient than conventional mixture model approaches. We demonstrate our method on the normal tissue samples and show that the clusters are associated with tissue type as well as age. Conclusion Our proposed recursively-partitioned mixture model is an effective and computationally efficient method for clustering DNA methylation data.

Published in BMC Bioinformatics

ISSN: 1471-2105 (Online)
Publisher: BMC
Country of publisher: United Kingdom
LCC subjects: Medicine: Medicine (General): Computer applications to medicine. Medical informatics; Science: Biology (General)
Website: http://www.biomedcentral.com/bmcbioinformatics/

About the journal