Affiliation:
1. Pontifícia Universidade Católica de Minas Gerais, Brazil
2. Universidade Federal de Minas Gerais, Brazil
Abstract
ABSTRACT Context: in recent years, cluster analysis has stimulated researchers to explore new ways to understand data behavior. The computational ease of this method and its ability to generate consistent outputs, even in small datasets, explain that to some extent. However, researchers are often mistaken in holding that clustering is a terrain in which anything goes. The literature shows the opposite: they must be careful, especially regarding the effect of outliers on cluster formation. Objective: in this tutorial paper, we contribute to this discussion by presenting four clustering techniques and their respective advantages and disadvantages in the treatment of outliers. Methods: for that, we worked from a managerial dataset and analyzed it using k-means, PAM, DBSCAN, and FCM techniques. Results: our analyzes indicate that researchers have distinct clustering techniques for dealing with outliers accordingly. Conclusion: we concluded that researchers need to have a more diversified repertoire of clustering techniques. After all, this would give them two relevant empirical alternatives: choose the most appropriate technique for their research objectives or adopt a multi-method approach.
Reference44 articles.
1. A gentle introduction to Stata;Acock A. C.,2014
2. Identifying and treating outliers in finance;Adams J.;Financial Management,2019
3. Data clustering: Algorithms and applications;Aggarwal C.,2014
4. Economics of strategy;Besanko D.,2016
5. Introduction to deep learning using R: A step-by-step guide to learning and implementing deep learning models using R;Beysolow T.,2017
Cited by
8 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献