The degradation of performance of a state-of-the-art skin image classifier when applied to patient-driven internet search

Seung Seog Han; Cristian Navarrete-Dechent; Konstantinos Liopyris; Myoung Shin Kim; Gyeong Hun Park; Sang Seok Woo; Juhyun Park; Jung Won Shin; Bo Ri Kim; Min Jae Kim; Francisca Donoso; Francisco Villanueva; Cristian Ramirez; Sung Eun Chang; Allan Halpern; Seong Hwan Kim; Jung-Im Na

doi:10.1038/s41598-022-20632-7

Scientific Reports (Sep 2022)

The degradation of performance of a state-of-the-art skin image classifier when applied to patient-driven internet search

Seung Seog Han,
Cristian Navarrete-Dechent,
Konstantinos Liopyris,
Myoung Shin Kim,
Gyeong Hun Park,
Sang Seok Woo,
Juhyun Park,
Jung Won Shin,
Bo Ri Kim,
Min Jae Kim,
Francisca Donoso,
Francisco Villanueva,
Cristian Ramirez,
Sung Eun Chang,
Allan Halpern,
Seong Hwan Kim,
Jung-Im Na

Affiliations

Seung Seog Han: Department of Dermatology, I Dermatology Clinic
Cristian Navarrete-Dechent: Department of Dermatology, School of Medicine, Pontificia Universidad Católica de Chile
Konstantinos Liopyris: Department of Dermatology, University of Athens, Andreas Syggros Hospital of Skin and Venereal Diseases
Myoung Shin Kim: Department of Dermatology, Sanggye Paik Hospital, Inje University College of Medicine
Gyeong Hun Park: Department of Dermatology, Dongtan Sacred Heart Hospital, Hallym University College of Medicine
Sang Seok Woo: Department of Plastic and Reconstructive Surgery, Kangnam Sacred Heart Hospital, Hallym University College of Medicine
Juhyun Park: Department of Dermatology, Seoul National University Bundang Hospital
Jung Won Shin: Department of Dermatology, Seoul National University Bundang Hospital
Bo Ri Kim: Department of Dermatology, Seoul National University Bundang Hospital
Min Jae Kim: Department of Dermatology, Seoul National University Bundang Hospital
Francisca Donoso: Department of Dermatology, School of Medicine, Pontificia Universidad Católica de Chile
Francisco Villanueva: Department of Dermatology, School of Medicine, Pontificia Universidad Católica de Chile
Cristian Ramirez: Department of Dermatology, School of Medicine, Pontificia Universidad Católica de Chile
Sung Eun Chang: Department of Dermatology, Asan Medical Center, Ulsan University College of Medicine
Allan Halpern: Dermatology Service, Department of Medicine, Memorial Sloan Kettering Cancer Center
Seong Hwan Kim: Department of Plastic and Reconstructive Surgery, Kangnam Sacred Heart Hospital, Hallym University College of Medicine
Jung-Im Na: Department of Dermatology, Seoul National University Bundang Hospital

DOI: https://doi.org/10.1038/s41598-022-20632-7
Journal volume & issue: Vol. 12, no. 1
pp. 1 – 9

Abstract

Read online

Abstract Model Dermatology ( https://modelderm.com ; Build2021) is a publicly testable neural network that can classify 184 skin disorders. We aimed to investigate whether our algorithm can classify clinical images of an Internet community along with tertiary care center datasets. Consecutive images from an Internet skin cancer community (‘RD’ dataset, 1,282 images posted between 25 January 2020 to 30 July 2021; https://reddit.com/r/melanoma ) were analyzed retrospectively, along with hospital datasets (Edinburgh dataset, 1,300 images; SNU dataset, 2,101 images; TeleDerm dataset, 340 consecutive images). The algorithm’s performance was equivalent to that of dermatologists in the curated clinical datasets (Edinburgh and SNU datasets). However, its performance deteriorated in the RD and TeleDerm datasets because of insufficient image quality and the presence of out-of-distribution disorders, respectively. For the RD dataset, the algorithm’s Top-1/3 accuracy (39.2%/67.2%) and AUC (0.800) were equivalent to that of general physicians (36.8%/52.9%). It was more accurate than that of the laypersons using random Internet searches (19.2%/24.4%). The Top-1/3 accuracy was affected by inadequate image quality (adequate = 43.2%/71.3% versus inadequate = 32.9%/60.8%), whereas participant performance did not deteriorate (adequate = 35.8%/52.7% vs. inadequate = 38.4%/53.3%). In this report, the algorithm performance was significantly affected by the change of the intended settings, which implies that AI algorithms at dermatologist-level, in-distribution setting, may not be able to show the same level of performance in with out-of-distribution settings.

Published in Scientific Reports

ISSN: 2045-2322 (Online)
Publisher: Nature Portfolio
Country of publisher: United Kingdom
LCC subjects: Medicine; Science
Website: https://www.nature.com/srep/

About the journal