ViMedNER: A Medical Named Entity Recognition Dataset for Vietnamese-Reference-Cited by-同舟云学术

ViMedNER: A Medical Named Entity Recognition Dataset for Vietnamese

Published:2024-07-11 Issue:4 Volume:11 Page:
ISSN:2410-0218
Container-title:EAI Endorsed Transactions on Industrial Networks and Intelligent Systems
language:
Short-container-title:EAI Endorsed Trans Ind Net Intel Syst

Author:

Duong Pham Van,Trinh Tien-Dat,Nguyen Minh-Tien,Vu Huy-The,Pham Minh Chuan,Tuan Tran Manh,Son Le Hoang

Abstract

Named entity recognition (NER) is one of the most important tasks in natural language processing, which identifies entity boundaries and classifies them into pre-defined categories. In literature, NER systems have been developed for various languages but limited works have been conducted for Vietnamese. This mainly comes from the limitation of available and high-quality annotated data, especially for specific domains such as medicine and healthcare. In this paper, we introduce a new medical NER dataset, named ViMedNER, for recognizing Vietnamese medical entities. Unlike existing works designed for common or too-specific entities, we focus on entity types that can be used in common diagnostic and treatment scenarios, including disease names, the symptoms of the diseases, the cause of the diseases, the diagnostic, and the treatment. These entities facilitate the diagnosis and treatment of doctors for common diseases. Our dataset is collected from four well-known Vietnamese websites that are professional in terms of drag selling and disease diagnostics and annotated by domain experts with high agreement scores. To create benchmark results, strong NER baselines based on pre-trained language models including PhoBERT, XLM-R, ViDeBERTa, ViPubMedDeBERTa, and ViHealthBERT are implemented and evaluated on the dataset. Experiment results show that the performance of XLM-R is consistently better than that of the other pre-trained language models. Furthermore, additional experiments are conducted to explore the behavior of the baselines and the characteristics of our dataset.

Funder

Bộ Giáo dục và Ðào tạo

Publisher

European Alliance for Innovation n.o.

Reference57 articles.

1. Angeli, G., Premkumar, M.J. and Manning, C.D. (2015) Leveraging linguistic structure for open domain information extraction. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 344-354.

2. Lample, G., Ballesteros, M., Subramanian, S., Kawakami, K. and Dyer, C. (2016) Neural architectures for named entity recognition. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 260-270.

3. Li, X., Feng, J., Meng, Y., Han, Q., Wu, F. and Li, J. (2020) A unified mrc framework for named entity recognition. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 5849- 5859.

4. Puccetti, G., Chiarello, F. and Fantoni, G. (2021) A simple and fast method for named entity context extraction from patents. Expert Systems with Applications 184 (2021): 115570 .

5. Sang, E., Kim, T. and Meulder, F.D. (2003) Introduction to the conll-2003 shared task: Language-independent named entity recognition. In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003.