Supervised Convex Clustering-Reference-Cited by-同舟云学术

Supervised Convex Clustering

Published:2023-03-23 Issue:4 Volume:79 Page:3846-3858
ISSN:0006-341X
Container-title:Biometrics
language:en
Short-container-title:

Author:

Wang Minjie¹^ORCID,Yao Tianyi²,Allen Genevera I.³

Affiliation:

1. School of Statistics, University of Minnesota , Minneapolis, Minnesota , USA

2. Department of Statistics, Rice University , Houston, Texas , USA

3. Departments of Electrical and Computer Engineering, Statistics, and Computer Science, Rice University and Jan and Dan Duncan Neurological Research Institute, Baylor College of Medicine , Houston, Texas , USA

Abstract

Abstract Clustering has long been a popular unsupervised learning approach to identify groups of similar objects and discover patterns from unlabeled data in many applications. Yet, coming up with meaningful interpretations of the estimated clusters has often been challenging precisely due to their unsupervised nature. Meanwhile, in many real-world scenarios, there are some noisy supervising auxiliary variables, for instance, subjective diagnostic opinions, that are related to the observed heterogeneity of the unlabeled data. By leveraging information from both supervising auxiliary variables and unlabeled data, we seek to uncover more scientifically interpretable group structures that may be hidden by completely unsupervised analyses. In this work, we propose and develop a new statistical pattern discovery method named supervised convex clustering (SCC) that borrows strength from both information sources and guides towards finding more interpretable patterns via a joint convex fusion penalty. We develop several extensions of SCC to integrate different types of supervising auxiliary variables, to adjust for additional covariates, and to find biclusters. We demonstrate the practical advantages of SCC through simulations and a case study on Alzheimer's disease genomics. Specifically, we discover new candidate genes as well as new subtypes of Alzheimer's disease that can potentially lead to better understanding of the underlying genetic mechanisms responsible for the observed heterogeneity of cognitive decline in older adults.

Funder

National Science Foundation

National Institutes of Health

Publisher

Oxford University Press (OUP)

Subject

Applied Mathematics,General Agricultural and Biological Sciences,General Immunology and Microbiology,General Biochemistry, Genetics and Molecular Biology,General Medicine,Statistics and Probability

Link

https://onlinelibrary.wiley.com/doi/pdf/10.1111/biom.13860

Reference30 articles.

1. K-means clustering based on Gower similarity coefficient: A comparative study;Ali,2013

2. Semi-supervised methods to predict patient survival from gene expression data;Bair;PLoS Biology,2004

3. Semi-supervised clustering by seeding;Basu,2002

4. Active semi-supervision for pairwise constrained clustering;Basu,2004

5. Religious orders study and rush memory and aging project;Bennett;Journal of Alzheimer's Disease,2018

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Understanding the Potential of Grouping Algorithms for Genetics Clustering;2024 2nd International Conference on Artificial Intelligence and Machine Learning Applications Theme: Healthcare and Internet of Things (AIMLA);2024-03-15