Comparative Analysis of Artificial Intelligence Virtual Assistant and Large Language Models in Post-Operative Care

Sahar Borna; Cesar A. Gomez-Cabello; Sophia M. Pressman; Syed Ali Haider; Ajai Sehgal; Bradley C. Leibovich; Dave Cole; Antonio Jorge Forte

doi:10.3390/ejihpe14050093

European Journal of Investigation in Health, Psychology and Education (May 2024)

Comparative Analysis of Artificial Intelligence Virtual Assistant and Large Language Models in Post-Operative Care

Sahar Borna,
Cesar A. Gomez-Cabello,
Sophia M. Pressman,
Syed Ali Haider,
Ajai Sehgal,
Bradley C. Leibovich,
Dave Cole,
Antonio Jorge Forte

Affiliations

Sahar Borna: Division of Plastic Surgery, Mayo Clinic, Jacksonville, FL 32224, USA
Cesar A. Gomez-Cabello: Division of Plastic Surgery, Mayo Clinic, Jacksonville, FL 32224, USA
Sophia M. Pressman: Division of Plastic Surgery, Mayo Clinic, Jacksonville, FL 32224, USA
Syed Ali Haider: Division of Plastic Surgery, Mayo Clinic, Jacksonville, FL 32224, USA
Ajai Sehgal: Center for Digital Health, Mayo Clinic, Rochester, MN 55905, USA
Bradley C. Leibovich: Center for Digital Health, Mayo Clinic, Rochester, MN 55905, USA
Dave Cole: Center for Digital Health, Mayo Clinic, Rochester, MN 55905, USA
Antonio Jorge Forte: Division of Plastic Surgery, Mayo Clinic, Jacksonville, FL 32224, USA

DOI: https://doi.org/10.3390/ejihpe14050093
Journal volume & issue: Vol. 14, no. 5
pp. 1413 – 1424

Abstract

Read online

In postoperative care, patient education and follow-up are pivotal for enhancing the quality of care and satisfaction. Artificial intelligence virtual assistants (AIVA) and large language models (LLMs) like Google BARD and ChatGPT-4 offer avenues for addressing patient queries using natural language processing (NLP) techniques. However, the accuracy and appropriateness of the information vary across these platforms, necessitating a comparative study to evaluate their efficacy in this domain. We conducted a study comparing AIVA (using Google Dialogflow) with ChatGPT-4 and Google BARD, assessing the accuracy, knowledge gap, and response appropriateness. AIVA demonstrated superior performance, with significantly higher accuracy (mean: 0.9) and lower knowledge gap (mean: 0.1) compared to BARD and ChatGPT-4. Additionally, AIVA’s responses received higher Likert scores for appropriateness. Our findings suggest that specialized AI tools like AIVA are more effective in delivering precise and contextually relevant information for postoperative care compared to general-purpose LLMs. While ChatGPT-4 shows promise, its performance varies, particularly in verbal interactions. This underscores the importance of tailored AI solutions in healthcare, where accuracy and clarity are paramount. Our study highlights the necessity for further research and the development of customized AI solutions to address specific medical contexts and improve patient outcomes.

Published in European Journal of Investigation in Health, Psychology and Education

ISSN: 2174-8144 (Print); 2254-9625 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Medicine: Public aspects of medicine; Philosophy. Psychology. Religion: Psychology
Website: https://www.mdpi.com/journal/ejihpe

About the journal

Abstract

Keywords