Extraction of Traditional Chinese Medicine Entity: Design of a Novel Span-Level Named Entity Recognition Method With Distant Supervision-Reference-Cited by-同舟云学术

Extraction of Traditional Chinese Medicine Entity: Design of a Novel Span-Level Named Entity Recognition Method With Distant Supervision

Published:2021-06-14 Issue:6 Volume:9 Page:e28219
ISSN:2291-9694
Container-title:JMIR Medical Informatics
language:en
Short-container-title:JMIR Med Inform

Author:

Jia Qi^ORCID,Zhang Dezheng^ORCID,Xu Haifeng^ORCID,Xie Yonghong^ORCID

Abstract

Background Traditional Chinese medicine (TCM) clinical records contain the symptoms of patients, diagnoses, and subsequent treatment of doctors. These records are important resources for research and analysis of TCM diagnosis knowledge. However, most of TCM clinical records are unstructured text. Therefore, a method to automatically extract medical entities from TCM clinical records is indispensable. Objective Training a medical entity extracting model needs a large number of annotated corpus. The cost of annotated corpus is very high and there is a lack of gold-standard data sets for supervised learning methods. Therefore, we utilized distantly supervised named entity recognition (NER) to respond to the challenge. Methods We propose a span-level distantly supervised NER approach to extract TCM medical entity. It utilizes the pretrained language model and a simple multilayer neural network as classifier to detect and classify entity. We also designed a negative sampling strategy for the span-level model. The strategy randomly selects negative samples in every epoch and filters the possible false-negative samples periodically. It reduces the bad influence from the false-negative samples. Results We compare our methods with other baseline methods to illustrate the effectiveness of our method on a gold-standard data set. The F1 score of our method is 77.34 and it remarkably outperforms the other baselines. Conclusions We developed a distantly supervised NER approach to extract medical entity from TCM clinical records. We estimated our approach on a TCM clinical record data set. Our experimental results indicate that the proposed approach achieves a better performance than other baselines.

Publisher

JMIR Publications Inc.

Subject

Health Information Management,Health Informatics

Reference20 articles.

1. Modern bioinformatics meets traditional Chinese medicine

2. Deep learning with word embeddings improves biomedical named entity recognition

3. Biomedical named entity recognition using deep neural networks with contextual information

4. An attention-based deep learning model for clinical named entity recognition of Chinese electronic medical records

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Development and Application of Traditional Chinese Medicine Using AI Machine Learning and Deep Learning Strategies;The American Journal of Chinese Medicine;2024-01

2. Chinese Named Entity Recognition Method for Domain-Specific Text;Tehnicki vjesnik - Technical Gazette;2023-12-15

3. RDRS: Represent Document-level Relation with Sentence-level Relation by Distant Supervision;2023 IEEE International Conference on Control, Electronics and Computer Technology (ICCECT);2023-04-28

4. A Multigranularity Text Driven Named Entity Recognition CGAN Model for Traditional Chinese Medicine Literatures;Computational Intelligence and Neuroscience;2022-09-24

5. Information Extraction from the Text Data on Traditional Chinese Medicine: A Review on Tasks, Challenges, and Methods from 2010 to 2021;Evidence-Based Complementary and Alternative Medicine;2022-05-13