A pre-trained BERT for Korean medical natural language processing-Reference-Cited by-同舟云学术

A pre-trained BERT for Korean medical natural language processing

Published:2022-08-16 Issue:1 Volume:12 Page:
ISSN:2045-2322
Container-title:Scientific Reports
language:en
Short-container-title:Sci Rep

Author:

Kim Yoojoong,Kim Jong-Ho,Lee Jeong Moon,Jang Moon Joung,Yum Yun Jin,Kim Seongtae,Shin Unsub,Kim Young-Min,Joo Hyung Joon,Song Sanghoun

Abstract

AbstractWith advances in deep learning and natural language processing (NLP), the analysis of medical texts is becoming increasingly important. Nonetheless, despite the importance of processing medical texts, no research on Korean medical-specific language models has been conducted. The Korean medical text is highly difficult to analyze because of the agglutinative characteristics of the language, as well as the complex terminologies in the medical domain. To solve this problem, we collected a Korean medical corpus and used it to train the language models. In this paper, we present a Korean medical language model based on deep learning NLP. The model was trained using the pre-training framework of BERT for the medical context based on a state-of-the-art Korean language model. The pre-trained model showed increased accuracies of 0.147 and 0.148 for the masked language model with next sentence prediction. In the intrinsic evaluation, the next sentence prediction accuracy improved by 0.258, which is a remarkable enhancement. In addition, the extrinsic evaluation of Korean medical semantic textual similarity data showed a 0.046 increase in the Pearson correlation, and the evaluation for the Korean medical named entity recognition showed a 0.053 increase in the F1-score.

Funder

National Research Foundation of Korea

Korea Health Industry Development Institute

Publisher

Springer Science and Business Media LLC

Subject

Multidisciplinary

Link

https://www.nature.com/articles/s41598-022-17806-8.pdf

Reference26 articles.

1. Zhang, Y., Chen, Q., Yang, Z., Lin, H. & Lu, Z. BioWordVec, improving biomedical word embeddings with subword information and MeSH. Sci. Data 6, 1–9 (2019).

2. Mikolov, T., Chen, K., Corrado, G. & Dean, J. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781 (2013).

3. Bojanowski, P., Grave, E., Joulin, A. & Mikolov, T. Enriching word vectors with subword information. Trans. Assoc. Comput. Linguist. 5, 135–146 (2017).

4. Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018).

5. Lan, Z. et al. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942 (2019).

Cited by 23 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Pre-trained language models in medicine: A survey;Artificial Intelligence in Medicine;2024-08

2. Transformer models in biomedicine;BMC Medical Informatics and Decision Making;2024-07-29

3. Assessing GPT-4’s Performance in Delivering Medical Advice: Comparative Analysis With Human Experts;JMIR Medical Education;2024-07-08

4. The Use of Clinical Language Models Pretrained on Institutional EHR Data for Downstream Tasks;2024 21st International Joint Conference on Computer Science and Software Engineering (JCSSE);2024-06-19

5. Key traits of top answerers on Korean Social Q&A platforms: insights into user performance and entrepreneurial potential;Humanities and Social Sciences Communications;2024-06-10