Empirical assessment of ChatGPT’s answering capabilities in natural science and engineering

Lukas Schulze Balhorn; Jana M. Weber; Stefan Buijsman; Julian R. Hildebrandt; Martina Ziefle; Artur M. Schweidtmann

doi:10.1038/s41598-024-54936-7

Scientific Reports (Feb 2024)

Empirical assessment of ChatGPT’s answering capabilities in natural science and engineering

Lukas Schulze Balhorn,
Jana M. Weber,
Stefan Buijsman,
Julian R. Hildebrandt,
Martina Ziefle,
Artur M. Schweidtmann

Affiliations

Lukas Schulze Balhorn: Delft University of Technology
Jana M. Weber: Delft University of Technology
Stefan Buijsman: Delft University of Technology
Julian R. Hildebrandt: Human-Computer Interaction Center, Chair of Communication Science, RWTH Aachen University
Martina Ziefle: Human-Computer Interaction Center, Chair of Communication Science, RWTH Aachen University
Artur M. Schweidtmann: Delft University of Technology

DOI: https://doi.org/10.1038/s41598-024-54936-7
Journal volume & issue: Vol. 14, no. 1
pp. 1 – 11

Abstract

Read online

Abstract ChatGPT is a powerful language model from OpenAI that is arguably able to comprehend and generate text. ChatGPT is expected to greatly impact society, research, and education. An essential step to understand ChatGPT’s expected impact is to study its domain-specific answering capabilities. Here, we perform a systematic empirical assessment of its abilities to answer questions across the natural science and engineering domains. We collected 594 questions on natural science and engineering topics from 198 faculty members across five faculties at Delft University of Technology. After collecting the answers from ChatGPT, the participants assessed the quality of the answers using a systematic scheme. Our results show that the answers from ChatGPT are, on average, perceived as “mostly correct”. Two major trends are that the rating of the ChatGPT answers significantly decreases (i) as the educational level of the question increases and (ii) as we evaluate skills beyond scientific knowledge, e.g., critical attitude.

Published in Scientific Reports

ISSN: 2045-2322 (Online)
Publisher: Nature Portfolio
Country of publisher: United Kingdom
LCC subjects: Medicine; Science
Website: https://www.nature.com/srep/

About the journal