Extraction of Meaningful Information from Unstructured Clinical Notes Using Web Scraping-Reference-Cited by-同舟云学术

Extraction of Meaningful Information from Unstructured Clinical Notes Using Web Scraping

Published:2022-09-14 Issue:03 Volume:32 Page:
ISSN:0218-1266
Container-title:Journal of Circuits, Systems and Computers
language:en
Short-container-title:J CIRCUIT SYST COMP

Author:

Varshini K. Sukanya¹^ORCID,Uthra R. Annie¹

Affiliation:

1. CINTEL, SRM Institute of Science and Technology, Kattankulathur, Chengelpet, Tamil Nadu 603203, India

Abstract

In the medical field, the clinical notes taken by the doctor, nurse, or medical practitioner are considered to be one of the most important medical documents. These documents hold information regarding the patient including the patient’s current condition, family history, disease, symptoms, medications, lab test reports, and other vital information. Despite these documents holding important information regarding the patients, they cannot be used as the data are unstructured. Organizing a huge amount of data without any mistakes is highly impossible for humans, so ignoring unstructured data is not advisable. Hence, to overcome this issue, the web scraping method is used to extract the clinical notes from the Medical Transcription (MT) samples which hold many transcripted clinical notes of various departments. In the proposed method, Natural Language Processing (NLP) is used to pre-process the data, and the variants of the Term Frequency-Inverse Document Frequency (TF-IDF)-based vector model are used for the feature selection, thus extracting the required data from the clinical notes. The performance measures including the accuracy, precision, recall and F1 score are used in the identification of disease, and the result obtained from the proposed system is compared with the best performing machine learning algorithms including the Logistic Regression, Multinomial Naive Bayes, Random Forest classifier and Linear SVC. The result obtained proves that the Random Forest Classifier obtained a higher accuracy of 90% when compared to the other algorithms.

Publisher

World Scientific Pub Co Pte Ltd

Subject

Electrical and Electronic Engineering,Hardware and Architecture,Media Technology

Link

https://www.worldscientific.com/doi/pdf/10.1142/S021812662350041X

Reference31 articles.

1. Machine Learning and Decision Support in Critical Care

2. Making Big Data Useful for Health Care: A Summary of the Inaugural MIT Critical Data Conference

3. Securing Critical Infrastructures: Deep-Learning-Based Threat Detection in IIoT