An Effective Analysis of Data Clustering using Distance-based K- Means Algorithm

Author:

Ramkumar P.,Kalamani P.,Valarmathi C.,Sheela Devi M.

Abstract

Abstract Real-world data sets are regularly provides different and complementary features of information in an unsupervised way. Different types of algorithms have been proposed recently in the genre of cluster analysis. It is arduous to the user to determine well in advance which algorithm would be the most suitable for a given dataset. Techniques with respect to graphs are provides excellent results for this task. However, the existing techniques are easily vulnerable to outliers and noises with limited idea about edges comprised in the tree to divide a dataset. Thus, in some fields, the necessacity for better clustering algorithms it uses robust and dynamic methods to improve and simplify the entire process of data clustering has become an important research field. In this paper, a novel distance-based clustering algorithm called the entropic distance based K-means clustering algorithm (EDBK) is proposed to eradicate the outliers in effective way. This algorithm depends on the entropic distance between attributes of data points and some basic mathematical statistics operations. In this work, experiments are carry out using UCI datasets showed that EDBK method which outperforms the existing methods such as Artificial Bee Colony (ABC), k-means.

Publisher

IOP Publishing

Subject

General Physics and Astronomy

Reference20 articles.

1. Anomaly detection model based on data stream clustering;Yin,2017

2. Incremental semi-supervised clustering ensemble for high dimensional data clustering;Yu;IEEE Transactions on Knowledge and Data Engineering,2016

3. A rapid hybrid clustering algorithm for large volumes of high dimensional data;Rathore;IEEE Transactions on Knowledge and Data Engineering,2019

4. Tune up fuzzy C-means for big data: some novel hybrid clustering algorithms based on initial selection and incremental clustering;Tien;International Journal of Fuzzy Systems 19.,2017

5. Cluster forest based fuzzy logic for massive data clustering;Lahmar

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3