Anomaly Detection Methods for Categorical Data-Reference-Cited by-同舟云学术

Anomaly Detection Methods for Categorical Data

Published:2019-05-31 Issue:2 Volume:52 Page:1-35
ISSN:0360-0300
Container-title:ACM Computing Surveys
language:en
Short-container-title:ACM Comput. Surv.

Author:

Taha Ayman¹,Hadi Ali S.²

Affiliation:

1. Faculty of Computers and Information, Cairo University, Giza, Egypt

2. American University in Cairo, Egypt, and Cornell University, Ithaca, NY, USA

Abstract

Anomaly detection has numerous applications in diverse fields. For example, it has been widely used for discovering network intrusions and malicious events. It has also been used in numerous other applications such as identifying medical malpractice or credit fraud. Detection of anomalies in quantitative data has received a considerable attention in the literature and has a venerable history. By contrast, and despite the widespread availability use of categorical data in practice, anomaly detection in categorical data has received relatively little attention as compared to quantitative data. This is because detection of anomalies in categorical data is a challenging problem. Some anomaly detection techniques depend on identifying a representative pattern then measuring distances between objects and this pattern. Objects that are far from this pattern are declared as anomalies. However, identifying patterns and measuring distances are not easy in categorical data compared with quantitative data. Fortunately, several papers focussing on the detection of anomalies in categorical data have been published in the recent literature. In this article, we provide a comprehensive review of the research on the anomaly detection problem in categorical data. Previous review articles focus on either the statistics literature or the machine learning and computer science literature. This review article combines both literatures. We review 36 methods for the detection of anomalies in categorical data in both literatures and classify them into 12 different categories based on the conceptual definition of anomalies they use. For each approach, we survey anomaly detection methods, and then show the similarities and differences among them. We emphasize two important issues, the number of parameters each method requires and its time complexity. The first issue is critical, because the performance of these methods are sensitive to the choice of these parameters. The time complexity is also very important in real applications especially in big data applications. We report the time complexity if it is reported by the authors of the methods. If it is not, then we derive it ourselves and report it in this article. In addition, we discuss the common problems and the future directions of the anomaly detection in categorical data.

Publisher

Association for Computing Machinery (ACM)

Subject

General Computer Science,Theoretical Computer Science

Link

https://dl.acm.org/doi/pdf/10.1145/3312739

Reference198 articles.

1. On the Vital Areas of Intrusion Detection Systems in Wireless Sensor Networks

2. Outlier detection techniques for localization in wireless sensor networks: A survey;Abukhalaf Hala;Int. J. Future Gen. Commun. Netw.,2015

3. Charu C. Aggarwal. 2017. Outlier Analysis 2nd ed. Springer Cham. Charu C. Aggarwal. 2017. Outlier Analysis 2nd ed. Springer Cham.

4. Outlier detection for high dimensional data

5. Outlier detection in graph streams

Cited by 66 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Recent advances in anomaly detection in Internet of Things: Status, challenges, and perspectives;Computer Science Review;2024-11

2. HEOD: Human-assisted Ensemble Outlier Detection for cybersecurity;Computers & Security;2024-11

3. Detecting Outliers in Context of Clustering Imbalanced Categorical Data;International Conference on Information Systems Development;2024-09-09

4. Machine Learning-Based Anomaly Detection on Seawater Temperature Data with Oversampling;Journal of Marine Science and Engineering;2024-05-12

5. Anomaly diagnosis of connected autonomous vehicles: A survey;Information Fusion;2024-05