ALE: automated label extraction from GEO metadata

Cory B. Giles; Chase A. Brown; Michael Ripperger; Zane Dennis; Xiavan Roopnarinesingh; Hunter Porter; Aleksandra Perz; Jonathan D. Wren

doi:10.1186/s12859-017-1888-1

BMC Bioinformatics (Dec 2017)

ALE: automated label extraction from GEO metadata

Cory B. Giles,
Chase A. Brown,
Michael Ripperger,
Zane Dennis,
Xiavan Roopnarinesingh,
Hunter Porter,
Aleksandra Perz,
Jonathan D. Wren

Affiliations

Cory B. Giles: Arthritis & Clinical Immunology Program, Oklahoma Medical Research Foundation
Chase A. Brown: Arthritis & Clinical Immunology Program, Oklahoma Medical Research Foundation
Michael Ripperger: Vanderbilt University
Zane Dennis: Department of Computer Science, Baylor University
Xiavan Roopnarinesingh: Arthritis & Clinical Immunology Program, Oklahoma Medical Research Foundation
Hunter Porter: Arthritis & Clinical Immunology Program, Oklahoma Medical Research Foundation
Aleksandra Perz: Arthritis & Clinical Immunology Program, Oklahoma Medical Research Foundation
Jonathan D. Wren: Arthritis & Clinical Immunology Program, Oklahoma Medical Research Foundation

DOI: https://doi.org/10.1186/s12859-017-1888-1
Journal volume & issue: Vol. 18, no. S14
pp. 7 – 16

Abstract

Read online

Abstract Background NCBI’s Gene Expression Omnibus (GEO) is a rich community resource containing millions of gene expression experiments from human, mouse, rat, and other model organisms. However, information about each experiment (metadata) is in the format of an open-ended, non-standardized textual description provided by the depositor. Thus, classification of experiments for meta-analysis by factors such as gender, age of the sample donor, and tissue of origin is not feasible without assigning labels to the experiments. Automated approaches are preferable for this, primarily because of the size and volume of the data to be processed, but also because it ensures standardization and consistency. While some of these labels can be extracted directly from the textual metadata, many of the data available do not contain explicit text informing the researcher about the age and gender of the subjects with the study. To bridge this gap, machine-learning methods can be trained to use the gene expression patterns associated with the text-derived labels to refine label-prediction confidence. Results Our analysis shows only 26% of metadata text contains information about gender and 21% about age. In order to ameliorate the lack of available labels for these data sets, we first extract labels from the textual metadata for each GEO RNA dataset and evaluate the performance against a gold standard of manually curated labels. We then use machine-learning methods to predict labels, based upon gene expression of the samples and compare this to the text-based method. Conclusion Here we present an automated method to extract labels for age, gender, and tissue from textual metadata and GEO data using both a heuristic approach as well as machine learning. We show the two methods together improve accuracy of label assignment to GEO samples.

Published in BMC Bioinformatics

ISSN: 1471-2105 (Online)
Publisher: BMC
Country of publisher: United Kingdom
LCC subjects: Medicine: Medicine (General): Computer applications to medicine. Medical informatics; Science: Biology (General)
Website: http://www.biomedcentral.com/bmcbioinformatics/

About the journal

Abstract

Keywords