Affiliation:
1. Department of Information Security, Seoul Women’s University, Nowon-gu, Seoul 01797, Republic of Korea
Abstract
Spam detection is an essential and unavoidable problem in today’s society. Most of the existing studies have used string-based detection methods with models and have been conducted on a single language, especially with English datasets. However, in the current global society, research on languages other than English is needed. String-based spam detection methods perform different preprocessing steps depending on language type due to differences in grammatical characteristics. Therefore, our study proposes a text-processing method and a string-imaging method. The CNN 2D visualization technology used in this paper can be applied to datasets of various languages by processing the data as images, so they can be equally applied to languages other than English. In this study, English and Korean spam data were used. As a result of this study, the string-based detection models of RNN, LSTM, and CNN 1D showed average accuracies of 0.9871, 0.9906, and 0.9912, respectively. On the other hand, the CNN 2D image-based detection model was confirmed to have an average accuracy of 0.9957. Through this study, we present a solution that shows that image-based processing is more effective than string-based processing for string data and that multilingual processing is possible based on the CNN 2D model.
Subject
Electrical and Electronic Engineering,Computer Networks and Communications,Hardware and Architecture,Signal Processing,Control and Systems Engineering
Reference36 articles.
1. (2022, April 20). How to Spot Scam Texts on Your Smartphone. Available online: https://www.aarp.org/money/scams-fraud/info-2021/texts-smartphone.html.
2. Cao, J., and Lai, C. (2020, January 7–11). A Bilingual Multi-type Spam Detection Model Based on M-BERT. Proceedings of the GLOBECOM 2020–2020 IEEE Global Communications Conference, Taipei, Taiwan.
3. Survey of review spam detection using machine learning techniques;Crawford;J. Big Data,2015
4. Luo, W. (2022, January 20–22). Research and Implementation of Text Topic Classification Based on Text CNN. Proceedings of the 2022 3rd International Conference on Computer Vision, Image and Deep Learning & International Conference on Computer Engineering and Applications (CVIDL & ICCEA), Changchun, China.
5. Comparative analysis of machine learning methods to detect fake news in an Urdu language corpus;Rafique;PeerJ Comput. Sci.,2022
Cited by
6 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献