Combinatorial K-Means Clustering as a Machine Learning Tool Applied to Diabetes Mellitus Type 2-Reference-Cited by-同舟云学术

Combinatorial K-Means Clustering as a Machine Learning Tool Applied to Diabetes Mellitus Type 2

Published:2021-02-17 Issue:4 Volume:18 Page:1919
ISSN:1660-4601
Container-title:International Journal of Environmental Research and Public Health
language:en
Short-container-title:IJERPH

Author:

Nedyalkova Miroslava^ORCID,Madurga Sergio^ORCID,Simeonov Vasil^ORCID

Abstract

A new original procedure based on k-means clustering is designed to find the most appropriate clinical variables able to efficiently separate into groups similar patients diagnosed with diabetes mellitus type 2 (DMT2) and underlying diseases (arterial hypertonia (AH), ischemic heart disease (CHD), diabetic polyneuropathy (DPNP), and diabetic microangiopathy (DMA)). Clustering is a machine learning tool for discovering structures in datasets. Clustering has been proven to be efficient for pattern recognition based on clinical records. The considered combinatorial k-means procedure explores all possible k-means clustering with a determined number of descriptors and groups. The predetermined conditions for the partitioning were as follows: every single group of patients included patients with DMT2 and one of the underlying diseases; each subgroup formed in such a way was subject to partitioning into three patterns (good health status, medium health status, and degenerated health status); optimal descriptors for each disease and groups. The selection of the best clustering is obtained through the parameter called global variance, defined as the sum of all variance values of all clinical variables of all the clusters. The best clinical parameters are found by minimizing this global variance. This methodology has to identify a set of variables that are assumed to separate each underlying disease efficiently in three different subgroups of patients. The hierarchical clustering obtained for these four underlying diseases could be used to build groups of patients with correlated clinical data. The proposed methodology gives surmised results from complex data based on a relationship with the health status of the group and draws a picture of the prediction rate of the ongoing health status.

Publisher

MDPI AG

Subject

Health, Toxicology and Mutagenesis,Public Health, Environmental and Occupational Health

Link

https://www.mdpi.com/1660-4601/18/4/1919/pdf

Reference14 articles.

1. Current Techniques for Diabetes Prediction: Review and Case Study

2. Classification of Diabetes Disease Using Support Vector Machine;Anuja;Int. J. Eng. Res. Appl.,2013

3. Application of Data Mining Methods and Techniques for Diabetes Diagnosis;Rajesh;Int. J. Eng. Innov. Technol.,2012

4. Diagnosis of Diabetes Using Classification Mining Techniques

5. A Prediction Technique in Data Mining for Diabetes Mellitus;Harleen;J. Manag. Sci. Technol.,2016

Cited by 29 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A new classification and laparoscopic treatment of extrahepatic choledochal cyst;Clinics and Research in Hepatology and Gastroenterology;2024-08

2. Breast cancer symptom profile longitudinal changes: data mining study;BMJ Supportive & Palliative Care;2024-06-25

3. Pre-processing techniques using a machine learning approach to improve model accuracy in estimating oil palm leaf chlorophyll from portable chlorophyll meter measurement;IOP Conference Series: Earth and Environmental Science;2024-02-01

4. Predictors of the Success of Yacht Charter in Andalusia from a Leading P2P Platform Using Machine Learning;Springer Proceedings in Business and Economics;2024

5. Temporal and geographic distribution of gut microbial enterotypes associated with host thermogenesis characteristics in plateau pikas;Microbiology Spectrum;2023-12-12