ChatGPT yields low accuracy in determining LI-RADS scores based on free-text and structured radiology reports in German language

Philipp Fervers; Robert Hahnfeldt; Jonathan Kottlors; Anton Wagner; David Maintz; Daniel Pinto dos Santos; Daniel Pinto dos Santos; Simon Lennartz; Thorsten Persigehl

doi:10.3389/fradi.2024.1390774

Frontiers in Radiology (Jul 2024)

ChatGPT yields low accuracy in determining LI-RADS scores based on free-text and structured radiology reports in German language

Philipp Fervers,
Robert Hahnfeldt,
Jonathan Kottlors,
Anton Wagner,
David Maintz,
Daniel Pinto dos Santos,
Daniel Pinto dos Santos,
Simon Lennartz,
Thorsten Persigehl

Affiliations

Philipp Fervers: Department of Diagnostic and Interventional Radiology, University Cologne, Faculty of Medicine and University Hospital Cologne, Cologne, Germany
Robert Hahnfeldt: Department of Diagnostic and Interventional Radiology, University Cologne, Faculty of Medicine and University Hospital Cologne, Cologne, Germany
Jonathan Kottlors: Department of Diagnostic and Interventional Radiology, University Cologne, Faculty of Medicine and University Hospital Cologne, Cologne, Germany
Anton Wagner: Department of Diagnostic and Interventional Radiology, University Cologne, Faculty of Medicine and University Hospital Cologne, Cologne, Germany
David Maintz: Department of Diagnostic and Interventional Radiology, University Cologne, Faculty of Medicine and University Hospital Cologne, Cologne, Germany
Daniel Pinto dos Santos: Department of Diagnostic and Interventional Radiology, University Cologne, Faculty of Medicine and University Hospital Cologne, Cologne, Germany
Daniel Pinto dos Santos: Department of Diagnostic and Interventional Radiology, Goethe University Frankfurt am Main, University Hospital Frankfurt, Frankfurt am Main, Germany
Simon Lennartz: Department of Diagnostic and Interventional Radiology, University Cologne, Faculty of Medicine and University Hospital Cologne, Cologne, Germany
Thorsten Persigehl: Department of Diagnostic and Interventional Radiology, University Cologne, Faculty of Medicine and University Hospital Cologne, Cologne, Germany

DOI: https://doi.org/10.3389/fradi.2024.1390774
Journal volume & issue: Vol. 4

Abstract

Read online

BackgroundTo investigate the feasibility of the large language model (LLM) ChatGPT for classifying liver lesions according to the Liver Imaging Reporting and Data System (LI-RADS) based on MRI reports, and to compare classification performance on structured vs. unstructured reports.MethodsLI-RADS classifiable liver lesions were included from German written structured and unstructured MRI reports with report of size, location, and arterial phase contrast enhancement as minimum inclusion requirements. The findings sections of the reports were propagated to ChatGPT (GPT-3.5), which was instructed to determine LI-RADS scores for each classifiable liver lesion. Ground truth was established by two radiologists in consensus. Agreement between ground truth and ChatGPT was assessed with Cohen's kappa. Test-retest reliability was assessed by passing a subset of n = 50 lesions five times to ChatGPT, using the intraclass correlation coefficient (ICC).Results205 MRIs from 150 patients were included. The accuracy of ChatGPT at determining LI-RADS categories was poor (53% and 44% on unstructured and structured reports). The agreement to the ground truth was higher (k = 0.51 and k = 0.44), the mean absolute error in LI-RADS scores was lower (0.5 ± 0.5 vs. 0.6 ± 0.7, p < 0.05), and the test-retest reliability was higher (ICC = 0.81 vs. 0.50), in free-text compared to structured reports, respectively, although structured reports comprised the minimum required imaging features significantly more frequently (Chi-square test, p < 0.05).ConclusionsChatGPT attained only low accuracy when asked to determine LI-RADS scores from liver imaging reports. The superior accuracy and consistency throughout free-text reports might relate to ChatGPT's training process.Clinical relevance statementOur study indicates both the necessity of optimization of LLMs for structured clinical data input and the potential of LLMs for creating machine-readable labels based on large free-text radiological databases.

Published in Frontiers in Radiology

ISSN: 2673-8740 (Online)
Publisher: Frontiers Media S.A.
Country of publisher: Switzerland
LCC subjects: Medicine: Medicine (General): Medical physics. Medical radiology. Nuclear medicine
Website: https://www.frontiersin.org/journals/radiology

About the journal

Abstract

Keywords