H-mrk-means: Enhanced Heuristic mrk-means for Linear Time Clustering of Big Data Using Hybrid Meta-heuristic Algorithm

Author:

Puri Digvijay1ORCID,Gupta Deepak1ORCID

Affiliation:

1. Department of CSE & IT, Jaypee University of Information Technology, Waknaghat, India

Abstract

Big data is generally derived with a large volume and combined categories of attributes like categorical and numerical. Among them, [Formula: see text]-prototypes have been adopted into MapReduce structure, and thus, it provides a better solution for the huge range of data. However, [Formula: see text]-prototypes need to compute all distances among every data point and cluster centres. Moreover, the computations of distances are redundant as data points are often present in similar clusters after fewer iterations. Nowadays, to cluster huge-scale datasets, one of the efficient solutions is [Formula: see text]-means. However, [Formula: see text]-means is not intrinsically appropriate to execute in MapReduce due to the iterative nature of this technique. Moreover, for every iteration, [Formula: see text]-means should perform an independent MapReduce job but, it leads to higher Input/Output (I/O) overhead at every iteration. This research paper presents a novel enhanced linear time clustering for handling big data called Heuristic mrk-means (H-mrk-means) using optimized [Formula: see text]-means on the MapReduce model. In order to manage big data that is time series in nature, the sampling and MapReduce framework are adopted, which utilize different machines for processing data. Before initiating the main clustering process, a sampling process is adopted to get the noteworthy information. The two main phases of the developed method are the map phase (divide and conquer) and the reduce phase (final clustering). In the map phase, the data are divided into diverse chunks that should be stored in assigned machines. In the reduce phase, data clustering is performed. Here, the cluster centroid of data is tuned with the help of hybrid Tunicate-Deer Hunting Optimization (T-DHO) algorithm by attaining a newly derived objective function. This type of optimal tuning of solution enhances the efficiency of clustering when compared over normal iterative [Formula: see text]-means and mrk-means clustering. The experimental evaluation on varied counts of chunks using the proposed H-mrk-means has attained higher quality of clustering results and faster execution times evaluated with other clustering approaches.

Publisher

World Scientific Pub Co Pte Ltd

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3