Machine learning models for prediction of double and triple burdens of non-communicable diseases in Bangladesh-Reference-Cited by-同舟云学术

Machine learning models for prediction of double and triple burdens of non-communicable diseases in Bangladesh

Published:2024-03-20 Issue:3 Volume:56 Page:426-444
ISSN:0021-9320
Container-title:Journal of Biosocial Science
language:en
Short-container-title:J. Biosoc. Sci.

Author:

Al-Zubayer Md. Akib^ORCID,Alam Khorshed^ORCID,Shanto Hasibul Hasan^ORCID,Maniruzzaman Md.,Majumder Uttam Kumar,Ahammed Benojir^ORCID

Abstract

AbstractIncreasing prevalence of non-communicable diseases (NCDs) has become the leading cause of death and disability in Bangladesh. Therefore, this study aimed to measure the prevalence of and risk factors for double and triple burden of NCDs (DBNCDs and TBNCDs), considering diabetes, hypertension, and overweight and obesity as well as establish a machine learning approach for predicting DBNCDs and TBNCDs. A total of 12,151 respondents from the 2017 to 2018 Bangladesh Demographic and Health Survey were included in this analysis, where 10%, 27.4%, and 24.3% of respondents had diabetes, hypertension, and overweight and obesity, respectively. Chi-square test and multilevel logistic regression (LR) analysis were applied to select factors associated with DBNCDs and TBNCDs. Furthermore, six classifiers including decision tree (DT), LR, naïve Bayes (NB), k-nearest neighbour (KNN), random forest (RF), and extreme gradient boosting (XGBoost) with three cross-validation protocols (K2, K5, and K10) were adopted to predict the status of DBNCDs and TBNCDs. The classification accuracy (ACC) and area under the curve (AUC) were computed for each protocol and repeated 10 times to make them more robust, and then the average ACC and AUC were computed. The prevalence of DBNCDs and TBNCDs was 14.3% and 2.3%, respectively. The findings of this study revealed that DBNCDs and TBNCDs were significantly influenced by age, sex, marital status, wealth index, education and geographic region. Compared to other classifiers, the RF-based classifier provides the highest ACC and AUC for both DBNCDs (ACC = 81.06% and AUC = 0.93) and TBNCDs (ACC = 88.61% and AUC = 0.97) for the K10 protocol. A combination of considered two-step factor selections and RF-based classifier can better predict the burden of NCDs. The findings of this study suggested that decision-makers might adopt suitable decisions to control and prevent the burden of NCDs using RF classifiers.

Publisher

Cambridge University Press (CUP)

Reference57 articles.

1. Predicting overweight and obesity in adulthood from body mass index values in childhood and adolescence;Guo;The American Journal of Clinical Nutrition,2002

2. Risk stratification for early detection of diabetes and hypertension in resource-limited settings: machine learning analysis;Boutilier;Journal of Medical Internet Research,2021

3. An Introduction to Statistical Learning

4. Bunkhumpornpat, C , Sinapiromsaran, K and Lursinsap, C (2011). MUTE: Majority under-sampling technique. In 2011 8th International Conference on Information, Communications & Signal Processing, IEEE, Singapore, pp. 1–4.

5. Montañez, CA , Fergus, P , Hussain, A , Al-Jumeily, D , Abdulaimma, B , Hind, J and Radi, N (2017). Machine learning approaches for the prediction of obesity using publicly available genetic profiles. In 2017 International Joint Conference on Neural Networks (IJCNN), IEEE, Anchorage, AK, USA, pp. 2743–2750.