Classification of cervical biopsy free-text diagnoses through linear-classifier based natural language processing

Jim Wei-Chun Hsu; Paul Christensen; Yimin Ge; S. Wesley Long

doi:10.1016/j.jpi.2022.100123

Journal of Pathology Informatics (Jan 2022)

Classification of cervical biopsy free-text diagnoses through linear-classifier based natural language processing

Jim Wei-Chun Hsu,
Paul Christensen,
Yimin Ge,
S. Wesley Long

Affiliations

Jim Wei-Chun Hsu: Department of Pathology and Genomic Medicine, Houston Methodist Hospital, Houston, Texas, USA
Paul Christensen: Department of Pathology and Genomic Medicine, Houston Methodist Hospital, Houston, Texas, USA; Department of Pathology and Laboratory Medicine, Weill Cornell Medical College, New York, New York, USA
Yimin Ge: Department of Pathology and Genomic Medicine, Houston Methodist Hospital, Houston, Texas, USA; Department of Pathology and Laboratory Medicine, Weill Cornell Medical College, New York, New York, USA
S. Wesley Long: Department of Pathology and Genomic Medicine, Houston Methodist Hospital, Houston, Texas, USA; Department of Pathology and Laboratory Medicine, Weill Cornell Medical College, New York, New York, USA; Houston Methodist Research Institute and Department of Pathology and Genomic Medicine, Houston Methodist Hospital, Houston, Texas, USA; Corresponding author at: Houston Methodist Hospital, 6565 Fannin St, Houston, TX 77004, USA.

DOI: https://doi.org/10.1016/j.jpi.2022.100123
Journal volume & issue: Vol. 13
p. 100123

Abstract

Read online

Routine cervical cancer screening has significantly decreased the incidence and mortality of cervical cancer. As selection of proper screening modalities depends on well-validated clinical decision algorithms, retrospective review correlating cytology and HPV test results with cervical biopsy diagnosis is essential for validating and revising these algorithms to changing technologies, demographics, and optimal clinical practices. However, manual categorization of the free-text biopsy diagnosis into discrete categories is extremely laborious due to the overwhelming number of specimens, which may lead to significant error and bias. Advances in machine learning and natural language processing (NLP), particularly over the last decade, have led to significant accomplishments and impressive performance in computer-based classification tasks. In this work, we apply an efficient version of an NLP framework, FastText™, to an annotated cervical biopsy dataset to create a supervised classifier that can assign accurate biopsy categories to free-text biopsy interpretations with high concordance to manually annotated data (>99.6%). We present cases where the machine-learning classifier disagrees with previous annotations and examine these discrepant cases after referee review by an expert pathologist. We also show that the classifier is robust on an untrained external dataset, achieving a concordance of 97.7%. In conclusion, we demonstrate a useful application of NLP to a real-world pathology classification task and highlight the benefits and limitations of this approach.

Published in Journal of Pathology Informatics

ISSN: 2229-5089 (Print); 2153-3539 (Online)
Publisher: Elsevier
Country of publisher: United States
LCC subjects: Medicine: Medicine (General): Computer applications to medicine. Medical informatics; Medicine: Pathology
Website: https://www.journals.elsevier.com/journal-of-pathology-informatics

About the journal

Abstract

Keywords