Lessons Learned from Data Mining of WHO Mortality Database-Reference-Cited by-同舟云学术

Lessons Learned from Data Mining of WHO Mortality Database

Published:2011 Issue:04 Volume:50 Page:380-385
ISSN:0026-1270
Container-title:Methods of Information in Medicine
language:en
Short-container-title:Methods Inf Med

Author:

Paoin W.

Abstract

SummaryObjectives: The objectives of this research were to test the ability of classification algorithms to predict the cause of death in the mortality data with unknown causes, to find association between common causes of death, to identify groups of countries based on their common causes of death, and to extract knowledge gained from data mining of the World Health Organization mortality database.Methods: The WEKA software version 3.5.3 was used for classification, clustering and association analysis of the World Health Organization mortality database which contained 1,109,537 records. Three major steps were performed: Step 1 – preprocessing of data to convert all records into suitable formats for each type of analysis algorithm; Step 2 – analyzing data using the C4.5 decision tree and Naïve Bayes classification algorithm, K-means clustering algorithm and Apriori association analysis algorithm; Step 3 – interpretation of results and hypothesis testing after clustering analysis.Results: Using a C4.5 decision tree classifier to predict cause of death, we obtained 440 leaf nodes that correctly classify death instances with an accuracy of 40.06%. Naïve Bayes classification algorithm calculated probability of death from each disease that correctly classify death instances with an accuracy of 28.13%. K means clustering divided the data into four clusters with 189, 59, 65, 144 country-years in each cluster. A Chi-square was used to test discriminate disease differences found in each cluster which had different diseases as predominant causes of death. Apriori association analysis produced association rules of linkage among cancer of the lung, hypertension and cerebrovascular diseases. These were found in the top five leading causes of death with 99–100% confidence level.Conclusion: Classification tools produced the poorest results in predicting cause of death. Given the inadequacy of variables in the WHO database, creation of a classification model to predict specific cause of death was impossible. Clustering and association tools yielded interesting results that could be used to identify new areas of interest in mortality data analysis. This can be used in data mining analysis to help solve some quality problems in mortality data.

Publisher

Georg Thieme Verlag KG

Subject

Health Information Management,Advanced and Specialized Nursing,Health Informatics

Link

http://www.thieme-connect.de/products/ejournals/pdf/10.3414/ME10-02-0019.pdf

Reference21 articles.

1. Han JW, Kamber M. Data Mining: Concepts and Techniques. 2nd ed. CA: Elsevier Inc; 2007. pp 5-27.

2. Tan PN, Steinbach M, Kumar V. Introduction to Data Mining. MA: Pearson Education Inc; 2006.

3. Biomedical Data Mining

4. Thailand Ministry of Public Health. Public Health Statistics, A.D. 1996-2005. 2005.

5. Infant and child mortality in developing countries: Analysing the data for Robust determinants

Cited by 8 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Mortality Prediction from Hospital-Acquired Infections in Trauma Patients Using an Unbalanced Dataset;Healthcare Informatics Research;2020-10-31

2. Predictive Apriori Algorithm in Youth Suicide Prevention by Screening Depressive Symptoms from Patient Health Questionnaire-9;TEM Journal;2019-11-30

3. Using Data Analytics to Predict Hospital Mortality in Sepsis Patients;International Journal of Healthcare Information Systems and Informatics;2019-07

4. Molecular Predicting Drought Tolerance in Maize Inbred Lines by Machine Learning Approaches;2019-03-16

5. Data Mining Approach in Mortality Projection: A Review Study;Advanced Science Letters;2018-03-01