Sparse Fuzzy C-Means Clustering with Lasso Penalty-Reference-Cited by-同舟云学术

Sparse Fuzzy C-Means Clustering with Lasso Penalty

Published:2024-09-13 Issue:9 Volume:16 Page:1208
ISSN:2073-8994
Container-title:Symmetry
language:en
Short-container-title:Symmetry

Author:

Parveen Shazia¹,Yang Miin-Shen¹^ORCID

Affiliation:

1. Department of Applied Mathematics, Chung Yuan Christian University, Taoyuan 32023, Taiwan

Abstract

Clustering is a technique of grouping data into a homogeneous structure according to the similarity or dissimilarity measures between objects. In clustering, the fuzzy c-means (FCM) algorithm is the best-known and most commonly used method and is a fuzzy extension of k-means in which FCM has been widely used in various fields. Although FCM is a good clustering algorithm, it only treats data points with feature components under equal importance and has drawbacks for handling high-dimensional data. The rapid development of social media and data acquisition techniques has led to advanced methods of collecting and processing larger, complex, and high-dimensional data. However, with high-dimensional data, the number of dimensions is typically immaterial or irrelevant. For features to be sparse, the Lasso penalty is capable of being applied to feature weights. A solution for FCM with sparsity is sparse FCM (S-FCM) clustering. In this paper, we propose a new S-FCM, called S-FCM-Lasso, which is a new type of S-FCM based on the Lasso penalty. The irrelevant features can be diminished towards exactly zero and assigned zero weights for unnecessary characteristics by the proposed S-FCM-Lasso. Based on various clustering performance measures, we compare S-FCM-Lasso with the S-FCM and other existing sparse clustering algorithms on several numerical and real-life datasets. Comparisons and experimental results demonstrate that, in terms of these performance measures, the proposed S-FCM-Lasso performs better than S-FCM and existing sparse clustering algorithms. This validates the efficiency and usefulness of the proposed S-FCM-Lasso algorithm for high-dimensional datasets with sparsity.

Funder

National Science and Technology Council, Taiwan

Publisher

MDPI AG

Link

https://www.mdpi.com/2073-8994/16/9/1208/pdf

Reference42 articles.

1. Model-based Gaussian and non-Gaussian clustering;Banfield;Biometrics,1993

2. On convergence and parameter selection of the EM and DA-EM algorithms for Gaussian mixtures;Yu;Pattern Recognit.,2018

3. A non-parametric approach to simplicity clustering;Hines;Appl. Artif. Intell.,2007

4. Adaptive nonparametric clustering;Efimov;IEEE Trans. Inf. Theory,2019

5. A comparative study of divisive and agglomerative hierarchical clustering algorithms;Roux;J. Classif.,2018