A Dataset for Evaluating Contextualized Representation of Biomedical Concepts in Language Models

Hossein Rouhizadeh; Irina Nikishina; Anthony Yazdani; Alban Bornet; Boya Zhang; Julien Ehrsam; Christophe Gaudet-Blavignac; Nona Naderi; Douglas Teodoro

doi:10.1038/s41597-024-03317-w

Scientific Data (May 2024)

A Dataset for Evaluating Contextualized Representation of Biomedical Concepts in Language Models

Hossein Rouhizadeh,
Irina Nikishina,
Anthony Yazdani,
Alban Bornet,
Boya Zhang,
Julien Ehrsam,
Christophe Gaudet-Blavignac,
Nona Naderi,
Douglas Teodoro

Affiliations

Hossein Rouhizadeh: Department of Radiology and Medical Informatics, Faculty of Medicine, University of Geneva
Irina Nikishina: Department of Informatics, University of Hamburg
Anthony Yazdani: Department of Radiology and Medical Informatics, Faculty of Medicine, University of Geneva
Alban Bornet: Department of Radiology and Medical Informatics, Faculty of Medicine, University of Geneva
Boya Zhang: Department of Radiology and Medical Informatics, Faculty of Medicine, University of Geneva
Julien Ehrsam: Department of Radiology and Medical Informatics, Faculty of Medicine, University of Geneva
Christophe Gaudet-Blavignac: Department of Radiology and Medical Informatics, Faculty of Medicine, University of Geneva
Nona Naderi: Laboratoire Interdisciplinaire des Sciences du Numerique, CNRS, Paris-Saclay University
Douglas Teodoro: Department of Radiology and Medical Informatics, Faculty of Medicine, University of Geneva

DOI: https://doi.org/10.1038/s41597-024-03317-w
Journal volume & issue: Vol. 11, no. 1
pp. 1 – 13

Abstract

Read online

Abstract Due to the complexity of the biomedical domain, the ability to capture semantically meaningful representations of terms in context is a long-standing challenge. Despite important progress in the past years, no evaluation benchmark has been developed to evaluate how well language models represent biomedical concepts according to their corresponding context. Inspired by the Word-in-Context (WiC) benchmark, in which word sense disambiguation is reformulated as a binary classification task, we propose a novel dataset, BioWiC, to evaluate the ability of language models to encode biomedical terms in context. BioWiC comprises 20’156 instances, covering over 7’400 unique biomedical terms, making it the largest WiC dataset in the biomedical domain. We evaluate BioWiC both intrinsically and extrinsically and show that it could be used as a reliable benchmark for evaluating context-dependent embeddings in biomedical corpora. In addition, we conduct several experiments using a variety of discriminative and generative large language models to establish robust baselines that can serve as a foundation for future research.

Published in Scientific Data

ISSN: 2052-4463 (Online)
Publisher: Nature Portfolio
Country of publisher: United Kingdom
LCC subjects: Science
Website: https://www.nature.com/sdata/

About the journal