Data-driven Approach to Age Prediction on Patients Diabetes and Cardiovascular Diseases Using Machine Learning: National Health and Nutrition Health Survey (Nhanes)

Author:

Abbas Irfan1

Affiliation:

1. Universitas Muhammadiyah Berau

Abstract

Abstract Background Diabetes and cardiovascular disease are two of the main causes of death in the United States. Identifying and predicting these diseases in patients is the first step towards stopping their progression. We evaluate the capabilities of machine learning models in detecting at-risk patients using survey data (and laboratory results), and identify key variables within the data contributing to these diseases among the patients. Methods Our research explores data-driven approaches which utilize supervised machine learning models to identify patients with such diseases. Using the National Health and Nutrition Examination Survey (NHANES) dataset, we conduct an exhaustive search of all available feature variables within the data to develop models for cardiovascular, prediabetes, and diabetes detection. Using different time-frames and feature sets for the data (based on laboratory data), multiple machine learning models (Support vector machines and adaptive boosting) were evaluated on their classification performance. The models were then combined to develop a weighted ensemble model, capable of leveraging the performance of the disparate models to improve detection accuracy. Information gain of tree-based models was used to identify the key variables within the patient data that contributed to the detection of at-risk patients in each of the diseases classes by the data-learned models. Results Diabetes and cardiovascular disease (CVD) are two of the leading causes of death in the United States. Detecting and predicting these diseases in patients is the first step to halting their progression. In this study, it was used Adaptive Boosting (AdaBoost) and Support Vector Machines (SVM) together as prediction. The purpose of this study was to knowing whether AdaBoost SVM could produce good accuracy. Tests were conducted using 50% data training and 50% data testing. Dot kernel were used to SVM. The highest accuracy value of AdaBoost SVM was accuracy 98.54%. Therefore it could be that AdaBoost can improve the performance of SVM in prediction of CVD desease severity Conclusion We conclude machine learned models based on survey questionnaire can provide an automated identification mechanism for patients at risk of diabetes and cardiovascular diseases. We also identify key contributors to the prediction, which can be further explored for their implications on electronic health records.

Publisher

Research Square Platform LLC

Reference48 articles.

1. Centers for disease control and prevention, “National Diabetes Statistics Report,” https://www.cdc.gov/diabetes/data/statistics-report/index.html.

2. A. Adler, “Using Machine Learning Techniques to Identify Key Risk Factors for Diabetes and Undiagnosed Diabetes,” May 2021, [Online]. Available: http://arxiv.org/abs/2105.09379

3. M. Niaz Imtiaz and A. Haque, “Predicting Type-2 Diabetes Using Machine Learning and Feature Selection Techniques.”

4. A. S. Abdalrada, J. Abawajy, T. Al-Quraishi, and S. M. S. Islam, “Machine learning models for prediction of co-occurrence of diabetes and cardiovascular diseases: a retrospective cohort study,” J Diabetes Metab Disord, vol. 21, no. 1, pp. 251–261, Jun. 2022, doi: 10.1007/s40200-021-00968-z.

5. A. Javaid et al., “Medicine 2032: The future of cardiovascular disease prevention with machine learning and digital health technology,” American Journal of Preventive Cardiology, vol. 12. Elsevier B.V., Dec. 01, 2022. doi: 10.1016/j.ajpc.2022.100379.

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3