BioWiC: An Evaluation Benchmark for Biomedical Concept Representation-Reference-Cited by-同舟云学术

BioWiC: An Evaluation Benchmark for Biomedical Concept Representation

Published:2023-11-11 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Rouhizadeh Hossein,Nikishina Irina,Yazdani Anthony,Bornet Alban^ORCID,Zhang Boya^ORCID,Ehrsam Julien,Gaudet-Blavignac Christophe,Naderi Nona,Teodoro Douglas^ORCID

Abstract

AbstractDue to the complexity of the biomedical domain, the ability to capture semantically meaningful representations of terms in context is a long-standing challenge. Despite important progress in the past years, no evaluation benchmark has been developed to evaluate how well language models represent biomedical concepts according to their corresponding context. Inspired by the Word-in-Context (WiC) benchmark, in which word sense disambiguation is reformulated as a binary classification task, we propose a novel dataset, BioWiC, to evaluate the ability of language models to encode biomedical terms in context. We evaluate BioWiC both intrinsically and extrinsically and show that it could be used as a reliable benchmark for evaluating context-dependent embeddings in biomedical corpora. In addition, we conduct several experiments using a variety of discriminative and generative large language models to establish robust baselines that can serve as a foundation for future research.

Publisher

Cold Spring Harbor Laboratory

Reference44 articles.

1. Detroja, K. , Bhensdadia, C. & Bhatt, B. S. A survey on relation extraction. Intell. Syst. with Appl. 200244 (2023).

2. Shi, J. et al. Knowledge-graph-enabled biomedical entity linking: a survey. World Wide Web 1–30 (2023).

3. French, E. & McInnes, B. T. An overview of biomedical entity linking throughout the years. J. Biomed. Informatic 104252 (2022).

4. Neural network-based approaches for biomedical relation classification: A review

5. Recent advances in biomedical literature mining