Impact of multi-source data augmentation on performance of convolutional neural networks for abnormality classification in mammography

InChan Hwang; Hari Trivedi; Beatrice Brown-Mulry; Linglin Zhang; Vineela Nalla; Aimilia Gastounioti; Judy Gichoya; Laleh Seyyed-Kalantari; Imon Banerjee; MinJae Woo

doi:10.3389/fradi.2023.1181190

Frontiers in Radiology (Jun 2023)

Impact of multi-source data augmentation on performance of convolutional neural networks for abnormality classification in mammography

InChan Hwang,
Hari Trivedi,
Beatrice Brown-Mulry,
Linglin Zhang,
Vineela Nalla,
Aimilia Gastounioti,
Judy Gichoya,
Laleh Seyyed-Kalantari,
Imon Banerjee,
MinJae Woo

Affiliations

InChan Hwang: School of Data Science and Analytics, Kennesaw State University, Kennesaw, GA, United States
Hari Trivedi: Department of Radiology, Emory University, Atlanta, GA, United States
Beatrice Brown-Mulry: School of Data Science and Analytics, Kennesaw State University, Kennesaw, GA, United States
Linglin Zhang: School of Data Science and Analytics, Kennesaw State University, Kennesaw, GA, United States
Vineela Nalla: Department of Information Technology, Kennesaw State University, Kennesaw, GA, United States
Aimilia Gastounioti: Mallinckrodt Institute of Radiology, Washington University in St. Louis, St. Louis, MO, United States
Judy Gichoya: Department of Radiology, Emory University, Atlanta, GA, United States
Laleh Seyyed-Kalantari: Department of Electrical Engineering and Computer Science, York University, Toronto, ON, Canada
Imon Banerjee: Department of Radiology, Mayo Clinic Arizona, Phoenix, AZ, United States
MinJae Woo: School of Data Science and Analytics, Kennesaw State University, Kennesaw, GA, United States

DOI: https://doi.org/10.3389/fradi.2023.1181190
Journal volume & issue: Vol. 3

Abstract

Read online

IntroductionTo date, most mammography-related AI models have been trained using either film or digital mammogram datasets with little overlap. We investigated whether or not combining film and digital mammography during training will help or hinder modern models designed for use on digital mammograms.MethodsTo this end, a total of six binary classifiers were trained for comparison. The first three classifiers were trained using images only from Emory Breast Imaging Dataset (EMBED) using ResNet50, ResNet101, and ResNet152 architectures. The next three classifiers were trained using images from EMBED, Curated Breast Imaging Subset of Digital Database for Screening Mammography (CBIS-DDSM), and Digital Database for Screening Mammography (DDSM) datasets. All six models were tested only on digital mammograms from EMBED.ResultsThe results showed that performance degradation to the customized ResNet models was statistically significant overall when EMBED dataset was augmented with CBIS-DDSM/DDSM. While the performance degradation was observed in all racial subgroups, some races are subject to more severe performance drop as compared to other races.DiscussionThe degradation may potentially be due to ( 1) a mismatch in features between film-based and digital mammograms ( 2) a mismatch in pathologic and radiological information. In conclusion, use of both film and digital mammography during training may hinder modern models designed for breast cancer screening. Caution is required when combining film-based and digital mammograms or when utilizing pathologic and radiological information simultaneously.

Published in Frontiers in Radiology

ISSN: 2673-8740 (Online)
Publisher: Frontiers Media S.A.
Country of publisher: Switzerland
LCC subjects: Medicine: Medicine (General): Medical physics. Medical radiology. Nuclear medicine
Website: https://www.frontiersin.org/journals/radiology

About the journal

Abstract

Keywords