Practical Web Spam Lifelong Machine Learning System with Automatic Adjustment to Current Lifecycle Phase-Reference-Cited by-同舟云学术

Practical Web Spam Lifelong Machine Learning System with Automatic Adjustment to Current Lifecycle Phase

Published:2019-02-20 Issue: Volume:2019 Page:1-16
ISSN:1939-0114
Container-title:Security and Communication Networks
language:en
Short-container-title:Security and Communication Networks

Author:

Luckner Marcin¹^ORCID

Affiliation:

1. Faculty of Mathematics and Information Science, Warsaw University of Technology, Koszykowa 75 Street, 00-662 Warsaw, Poland

Abstract

Machine learning techniques are a standard approach in spam detection. Their quality depends on the quality of the learning set, and when the set is out of date, the quality of classification falls rapidly. The most popular public web spam dataset that can be used to train a spam detector—WEBSPAM-UK2007—is over ten years old. Therefore, there is a place for a lifelong machine learning system that can replace the detectors based on a static learning set. In this paper, we propose a novel web spam recognition system. The system automatically rebuilds the learning set to avoid classification based on outdated data. Using a built-in automatic selection of the active classifier the system very quickly attains productive accuracy despite a limited learning set. Moreover, the system automatically rebuilds the learning set using external data from spam traps and popular web services. A test on real data from Quora, Reddit, and Stack Overflow proved the high recognition quality. Both the obtained average accuracy and the F-measure were 0.98 and 0.96 for semiautomatic and full–automatic mode, respectively.

Funder

National Science Center

Publisher

Hindawi Limited

Subject

Computer Networks and Communications,Information Systems

Link

http://downloads.hindawi.com/journals/scn/2019/6587020.pdf

Reference26 articles.

1. The WEKA data mining software

2. Efficient and effective spam filtering and re-ranking for large web datasets

3. Link-based web spam detection using weight properties

4. Stable web spam detection using features based on lexical items

5. Web Spam Detection: New Classification Features Based on Qualified Link Analysis and Language Models

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Android malware concept drift using system calls: Detection, characterization and challenges;Expert Systems with Applications;2022-11

2. GT2FS-SMOTE: An Intelligent Oversampling Approach Based Upon General Type-2 Fuzzy Sets to Detect Web Spam;Arabian Journal for Science and Engineering;2020-10-15