Accuracy, appropriateness, and readability of ChatGPT-4 and ChatGPT-3.5 in answering pediatric emergency medicine post-discharge questions

Mitul Gupta; Aiza Kahlun; Ria Sur; Pramiti Gupta; Andrew Kienstra; Winnie Whitaker; Graham Aufricht

doi:10.22470/pemj.2024.01074

Pediatric Emergency Medicine Journal (Apr 2025)

Accuracy, appropriateness, and readability of ChatGPT-4 and ChatGPT-3.5 in answering pediatric emergency medicine post-discharge questions

Mitul Gupta,
Aiza Kahlun,
Ria Sur,
Pramiti Gupta,
Andrew Kienstra,
Winnie Whitaker,
Graham Aufricht

Affiliations

Mitul Gupta: Department of Diagnostic Medicine, Dell Medical School, The University of Texas at Austin, Austin, TX, USA
Aiza Kahlun: Department of Diagnostic Medicine, Dell Medical School, The University of Texas at Austin, Austin, TX, USA
Ria Sur: Department of Diagnostic Medicine, Dell Medical School, The University of Texas at Austin, Austin, TX, USA
Pramiti Gupta: Undergraduate Program, The University of Texas at Austin, Austin, TX, USA
Andrew Kienstra: Department of Pediatric Emergency Medicine, Dell Medical School, The University of Texas at Austin, Austin, TX, USA
Winnie Whitaker: Department of Pediatric Emergency Medicine, Dell Medical School, The University of Texas at Austin, Austin, TX, USA
Graham Aufricht: Department of Pediatric Emergency Medicine, Dell Medical School, The University of Texas at Austin, Austin, TX, USA

DOI: https://doi.org/10.22470/pemj.2024.01074
Journal volume & issue: Vol. 12, no. 2
pp. 62 – 72

Abstract

Read online

Purpose Large language models (LLMs) like ChatGPT (OpenAI) are increasingly used in healthcare, raising questions about their accuracy and reliability for medical information. This study compared 2 versions of ChatGPT in answering post-discharge follow-up questions in the area of pediatric emergency medicine (PEM). Methods Twenty-three common post-discharge questions were posed to ChatGPT-4 and -3.5, with responses generated before and after a simplification request. Two blinded PEM physicians evaluated appropriateness and accuracy as the primary endpoint. Secondary endpoints included word count and readability. Six established reading scales were averaged, including the Automated Readability Index, Gunning Fog Index, Flesch-Kincaid Grade Level, Coleman-Liau Index, Simple Measure of Gobbledygook Grade Level, and Flesch Reading Ease. T-tests and Cohen’s kappa were used to determine differences and inter-rater agreement, respectively. Results The physician evaluations showed high appropriateness for both defaults (ChatGPT-4, 91.3%-100% vs. ChatGPT-3.5, 91.3%) and simplified responses (both 87.0%-91.3%). The accuracy was also high for default (87.0%-95.7% vs. 87.0%-91.3%) and simplified responses (both 82.6%-91.3%). The inter-rater agreement was fair overall (κ = 0.37; P < 0.001). For default responses, ChatGPT-4 produced longer outputs than ChatGPT-3.5 (233.0 ± 97.1 vs. 199.6 ± 94.7 words; P = 0.043), with a similar readability (13.3 ± 1.9 vs. 13.5 ± 1.8; P = 0.404). After simplification, both LLMs improved word count and readability (P < 0.001), with ChatGPT-4 achieving a readability suitable for the eighth grade students in the United States (7.7 ± 1.3 vs. 8.2 ± 1.5; P = 0.027). Conclusion The responses of ChatGPT-4 and -3.5 to post-discharge questions were deemed appropriate and accurate by the PEM physicians. While ChatGPT-4 showed an edge in simplifying language, neither LLM consistently met the recommended reading level of sixth grade students. These findings suggest a potential for LLMs to communicate with guardians.

Published in Pediatric Emergency Medicine Journal

ISSN: 2383-4897 (Print); 2508-5506 (Online)
Publisher: Korean Society of Pediatric Emergency Medicine
Country of publisher: Korea, Republic of
LCC subjects: Medicine
Website: https://www.pemj.org/

About the journal

Abstract

Keywords