Are large language models valid tools for patient information on lumbar disc herniation? The spine surgeons' perspective

Siegmund Lang; Jacopo Vitale; Tamás F. Fekete; Daniel Haschtmann; Raluca Reitmeir; Mario Ropelato; Jani Puhakka; Fabio Galbusera; Markus Loibl

Brain and Spine (Jan 2024)

Are large language models valid tools for patient information on lumbar disc herniation? The spine surgeons' perspective

Siegmund Lang,
Jacopo Vitale,
Tamás F. Fekete,
Daniel Haschtmann,
Raluca Reitmeir,
Mario Ropelato,
Jani Puhakka,
Fabio Galbusera,
Markus Loibl

Affiliations

Siegmund Lang: Department of Trauma Surgery, University Hospital Regensburg, Regensburg, Germany; Corresponding author. Department of Trauma Surgery, University Hospital Regensburg, Franz-Joseph-Strauss-Allee 11, 93053, Regensburg, Germany.
Jacopo Vitale: Spine Center, Schulthess Klinik, Zurich, Switzerland
Tamás F. Fekete: Spine Center, Schulthess Klinik, Zurich, Switzerland
Daniel Haschtmann: Spine Center, Schulthess Klinik, Zurich, Switzerland
Raluca Reitmeir: Spine Center, Schulthess Klinik, Zurich, Switzerland
Mario Ropelato: Spine Center, Schulthess Klinik, Zurich, Switzerland
Jani Puhakka: Spine Center, Schulthess Klinik, Zurich, Switzerland
Fabio Galbusera: Spine Center, Schulthess Klinik, Zurich, Switzerland
Markus Loibl: Spine Center, Schulthess Klinik, Zurich, Switzerland

Journal volume & issue: Vol. 4
p. 102804

Abstract

Read online

Introduction: Generative AI is revolutionizing patient education in healthcare, particularly through chatbots that offer personalized, clear medical information. Reliability and accuracy are vital in AI-driven patient education. Research question: How effective are Large Language Models (LLM), such as ChatGPT and Google Bard, in delivering accurate and understandable patient education on lumbar disc herniation? Material and methods: Ten Frequently Asked Questions about lumbar disc herniation were selected from 133 questions and were submitted to three LLMs. Six experienced spine surgeons rated the responses on a scale from “excellent” to “unsatisfactory,” and evaluated the answers for exhaustiveness, clarity, empathy, and length. Statistical analysis involved Fleiss Kappa, Chi-square, and Friedman tests. Results: Out of the responses, 27.2% were excellent, 43.9% satisfactory with minimal clarification, 18.3% satisfactory with moderate clarification, and 10.6% unsatisfactory. There were no significant differences in overall ratings among the LLMs (p = 0.90); however, inter-rater reliability was not achieved, and large differences among raters were detected in the distribution of answer frequencies. Overall, ratings varied among the 10 answers (p = 0.043). The average ratings for exhaustiveness, clarity, empathy, and length were above 3.5/5. Discussion and conclusion: LLMs show potential in patient education for lumbar spine surgery, with generally positive feedback from evaluators. The new EU AI Act, enforcing strict regulation on AI systems, highlights the need for rigorous oversight in medical contexts. In the current study, the variability in evaluations and occasional inaccuracies underline the need for continuous improvement. Future research should involve more advanced models to enhance patient-physician communication.

Published in Brain and Spine

ISSN: 2772-5294 (Online)
Publisher: Elsevier
Country of publisher: Netherlands
LCC subjects: Medicine: Internal medicine: Neurosciences. Biological psychiatry. Neuropsychiatry: Neurology. Diseases of the nervous system
Website: https://www.journals.elsevier.com/brain-and-spine

About the journal

Abstract

Keywords