Large language models for generating medical examinations: systematic review

Yaara Artsi; Vera Sorin; Eli Konen; Benjamin S. Glicksberg; Girish Nadkarni; Eyal Klang

doi:10.1186/s12909-024-05239-y

BMC Medical Education (Mar 2024)

Large language models for generating medical examinations: systematic review

Yaara Artsi,
Vera Sorin,
Eli Konen,
Benjamin S. Glicksberg,
Girish Nadkarni,
Eyal Klang

Affiliations

Yaara Artsi: Azrieli Faculty of Medicine, Bar-Ilan University
Vera Sorin: Department of Diagnostic Imaging, Chaim Sheba Medical Center
Eli Konen: Department of Diagnostic Imaging, Chaim Sheba Medical Center
Benjamin S. Glicksberg: Division of Data-Driven and Digital Medicine (D3M), Icahn School of Medicine at Mount Sinai
Girish Nadkarni: Division of Data-Driven and Digital Medicine (D3M), Icahn School of Medicine at Mount Sinai
Eyal Klang: Division of Data-Driven and Digital Medicine (D3M), Icahn School of Medicine at Mount Sinai

DOI: https://doi.org/10.1186/s12909-024-05239-y
Journal volume & issue: Vol. 24, no. 1
pp. 1 – 11

Abstract

Read online

Abstract Background Writing multiple choice questions (MCQs) for the purpose of medical exams is challenging. It requires extensive medical knowledge, time and effort from medical educators. This systematic review focuses on the application of large language models (LLMs) in generating medical MCQs. Methods The authors searched for studies published up to November 2023. Search terms focused on LLMs generated MCQs for medical examinations. Non-English, out of year range and studies not focusing on AI generated multiple-choice questions were excluded. MEDLINE was used as a search database. Risk of bias was evaluated using a tailored QUADAS-2 tool. Results Overall, eight studies published between April 2023 and October 2023 were included. Six studies used Chat-GPT 3.5, while two employed GPT 4. Five studies showed that LLMs can produce competent questions valid for medical exams. Three studies used LLMs to write medical questions but did not evaluate the validity of the questions. One study conducted a comparative analysis of different models. One other study compared LLM-generated questions with those written by humans. All studies presented faulty questions that were deemed inappropriate for medical exams. Some questions required additional modifications in order to qualify. Conclusions LLMs can be used to write MCQs for medical examinations. However, their limitations cannot be ignored. Further study in this field is essential and more conclusive evidence is needed. Until then, LLMs may serve as a supplementary tool for writing medical examinations. 2 studies were at high risk of bias. The study followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines.

Published in BMC Medical Education

ISSN: 1472-6920 (Online)
Publisher: BMC
Country of publisher: United Kingdom
LCC subjects: Education: Special aspects of education; Medicine
Website: https://bmcmededuc.biomedcentral.com

About the journal

Abstract

Keywords