A New Data Representation Based on Training Data Characteristics to Extract Drug Name Entity in Medical Text-Reference-Cited by-同舟云学术

A New Data Representation Based on Training Data Characteristics to Extract Drug Name Entity in Medical Text

Published:2016 Issue: Volume:2016 Page:1-16
ISSN:1687-5265
Container-title:Computational Intelligence and Neuroscience
language:en
Short-container-title:Computational Intelligence and Neuroscience

Author:

Sadikin Mujiono¹²^ORCID,Fanany Mohamad Ivan²^ORCID,Basaruddin T.²

Affiliation:

1. Faculty of Computer Science, Universitas Mercu Buana, l. Meruya Selatan No. 1, Kembangan, Jakarta Barat 11650, Indonesia

2. Machine Learning and Computer Vision Laboratory, Faculty of Computer Science, Universitas Indonesia, Depok, West Java 16424, Indonesia

Abstract

One essential task in information extraction from the medical corpus is drug name recognition. Compared with text sources come from other domains, the medical text mining poses more challenges, for example, more unstructured text, the fast growing of new terms addition, a wide range of name variation for the same drug, the lack of labeled dataset sources and external knowledge, and the multiple token representations for a single drug name. Although many approaches have been proposed to overwhelm the task, some problems remained with poor F-score performance (less than 0.75). This paper presents a new treatment in data representation techniques to overcome some of those challenges. We propose three data representation techniques based on the characteristics of word distribution and word similarities as a result of word embedding training. The first technique is evaluated with the standard NN model, that is, MLP. The second technique involves two deep network classifiers, that is, DBN and SAE. The third technique represents the sentence as a sequence that is evaluated with a recurrent NN model, that is, LSTM. In extracting the drug name entities, the third technique gives the best F-score performance compared to the state of the art, with its average F-score being 0.8645.

Funder

Higher Education Science and Technology Development

Publisher

Hindawi Limited

Subject

General Mathematics,General Medicine,General Neuroscience,General Computer Science

Link

http://downloads.hindawi.com/journals/cin/2016/3483528.pdf

Reference20 articles.

1. Unsupervised biomedical named entity recognition: Experiments with clinical and biological texts

2. Boosting drug named entity recognition using an aggregate classifier

3. The DDI corpus: An annotated corpus with pharmacological substances and drug–drug interactions

4. Mining Adverse Drug Reactions from online healthcare forums using Hidden Markov Model

5. Drug name recognition and classification in biomedical texts

Cited by 11 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Evaluation of Machine Learning Approach for Sentiment Analysis using Yelp Dataset;European Journal of Electrical Engineering and Computer Science;2023-12-10

2. Hybrid Deep Learning for Medication-Related Information Extraction From Clinical Texts in French: MedExt Algorithm Development Study;JMIR Medical Informatics;2021-03-16

3. Users' Emotions Analysis based on Hybrid Feature Extraction Techniques;International Journal of Scientific Research in Computer Science, Engineering and Information Technology;2020-12-10

4. Developing an Expert System Application to Detect Childs' Lung Disease;International Journal of Scientific Research in Computer Science, Engineering and Information Technology;2020-12-10

5. Hybrid Deep Learning for Medication-Related Information Extraction From Clinical Texts in French: MedExt Algorithm Development Study (Preprint);2020-01-23