Anomaly Detection in Log Files Using Selected Natural Language Processing Methods-Reference-Cited by-同舟云学术

Anomaly Detection in Log Files Using Selected Natural Language Processing Methods

Published:2022-05-18 Issue:10 Volume:12 Page:5089
ISSN:2076-3417
Container-title:Applied Sciences
language:en
Short-container-title:Applied Sciences

Author:

Ryciak Piotr^ORCID,Wasielewska Katarzyna^ORCID,Janicki Artur^ORCID

Abstract

In this article, we address the problem of detecting anomalies in system log files. Computer systems generate huge numbers of events, which are noted in event log files. While most of them report normal actions, an unusual entry may inform about a failure or malware infection. A human operator may easily miss such an entry; therefore, anomaly detection methods are used for this purpose. In our work, we used an approach known from the natural language processing (NLP) domain, which operates on so-called embeddings, that is vector representations of words or phrases. We describe an improved version of the LogEvent2Vec algorithm, proposed in 2020. In contrast to the original version, we propose a significant shortening of the analysis window, which both increased the accuracy of anomaly detection and made further analysis of suspicious sequences much easier. We experimented with various binary classifiers, such as decision trees or multilayer perceptrons (MLPs), and the Blue Gene/L dataset. We showed that selecting an optimal classifier (in this case, MLP) and a short log sequence gave very good results. The improved version of the algorithm yielded the best F1-score of 0.997, compared to 0.886 in the original version of the algorithm.

Funder

European Commission

Publisher

MDPI AG

Subject

Fluid Flow and Transfer Processes,Computer Science Applications,Process Chemistry and Technology,General Engineering,Instrumentation,General Materials Science

Link

https://www.mdpi.com/2076-3417/12/10/5089/pdf

Reference47 articles.

1. Detecting large-scale system problems by mining console logs

2. Advances and challenges in log analysis

3. On Vulnerability and Security Log analysis

4. A Survey on Automated Log Analysis for Reliability Engineering

5. Collecting router information for error diagnosis and troubleshooting in home networks

Cited by 12 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Interpretable Feature Learning in Multivariate Big Data Analysis for Network Monitoring;IEEE Transactions on Network and Service Management;2024-06

2. Fake news detection models using the largest social media ground-truth dataset (TruthSeeker);International Journal of Speech Technology;2024-06

3. Unsupervised Anomaly Detection in Sequential Process Data;Zeitschrift für Psychologie;2024-04

4. Landscape and Taxonomy of Online Parser-Supported Log Anomaly Detection Methods;IEEE Access;2024

5. A Comprehensive Review on Transforming Security and Privacy with NLP;Lecture Notes in Networks and Systems;2024