Automatic Classification of Cancer Pathology Reports: A Systematic Review

Thiago Santos; Amara Tariq; Judy Wawira Gichoya; Hari Trivedi; Imon Banerjee

doi:10.1016/j.jpi.2022.100003

Journal of Pathology Informatics (Jan 2022)

Automatic Classification of Cancer Pathology Reports: A Systematic Review

Thiago Santos,
Amara Tariq,
Judy Wawira Gichoya,
Hari Trivedi,
Imon Banerjee

Affiliations

Thiago Santos: Department of Computer Science, Emory University, Atlanta, GA, USA; Department of Biomedical Informatics, Emory School of Medicine, Atlanta, GA, USA; Corresponding author.
Amara Tariq: Department of Radiology, Mayo Clinic, Phoenix, AZ, USA
Judy Wawira Gichoya: Department of Biomedical Informatics, Emory School of Medicine, Atlanta, GA, USA; Department of Radiology, Emory School of Medicine, Atlanta, GA, USA
Hari Trivedi: Department of Biomedical Informatics, Emory School of Medicine, Atlanta, GA, USA; Department of Radiology, Emory School of Medicine, Atlanta, GA, USA
Imon Banerjee: Department of Radiology, Mayo Clinic, Phoenix, AZ, USA; Department of Computer Engineering, Arizona State University, AZ, USA

DOI: https://doi.org/10.1016/j.jpi.2022.100003
Journal volume & issue: Vol. 13
p. 100003

Abstract

Read online

Pathology reports primarily consist of unstructured free text and thus the clinical information contained in the reports is not trivial to access or query. Multiple natural language processing (NLP) techniques have been proposed to automate the coding of pathology reports via text classification. In this systematic review, we follow the guidelines proposed by the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA; Page et al., 2020: BMJ.) to identify the NLP systems for classifying pathology reports published between the years of 2010 and 2021. Based on our search criteria, a total of 3445 records were retrieved, and 25 articles met the final review criteria. We benchmarked the systems based on methodology, complexity of the prediction task and core types of NLP models: i) Rule-based and Intelligent systems, ii) statistical machine learning, and iii) deep learning. While certain tasks are well addressed by these models, many others have limitations and remain as open challenges, such as, extraction of many cancer characteristics (size, shape, type of cancer, others) from pathology reports. We investigated the final set of papers (25) and addressed their potential as well as their limitations. We hope that this systematic review helps researchers prioritize the development of innovated approaches to tackle the current limitations and help the advancement of cancer research.

Published in Journal of Pathology Informatics

ISSN: 2229-5089 (Print); 2153-3539 (Online)
Publisher: Elsevier
Country of publisher: United States
LCC subjects: Medicine: Medicine (General): Computer applications to medicine. Medical informatics; Medicine: Pathology
Website: https://www.journals.elsevier.com/journal-of-pathology-informatics

About the journal