SpotSpam: Intention Analysis–driven SMS Spam Detection Using BERT Embeddings-Reference-Cited by-同舟云学术

SpotSpam: Intention Analysis–driven SMS Spam Detection Using BERT Embeddings

Published:2022-08-31 Issue:3 Volume:16 Page:1-27
ISSN:1559-1131
Container-title:ACM Transactions on the Web
language:en
Short-container-title:ACM Trans. Web

Author:

Oswald C.¹^ORCID,Simon Sona Elza²^ORCID,Bhattacharya Arnab¹^ORCID

Affiliation:

1. Indian Institute of Technology, Kanpur, India

2. Indian Institute of Information Technology Design and Manufacturing, Kancheepuram, India

Abstract

Short Message Service (SMS) is one of the widely used mobile applications for global communication for personal and business purposes. Its widespread use for customer interaction, business updates, and reminders has made it a billion-dollar industry in “Text Marketing.” Along with valid SMS, a tsunami of spam messages also pop up that serve various purposes for the sender and the majority of them are fraudulent. Filtering spam SMS in an accurate manner is a crucial and challenging task that will benefit human lives both mentally and economically. Some of the challenges in the filtering of spam SMS include less number of characters, texts in informal languages, lack of public SMS spam corpus, and so on. Focusing solely on the textual features of the SMS is a major handicap of the existing methods, as it lacks in dynamically adapting to the increasing number of new keywords and jargon. In this article, we develop an intention-based approach of SMS spam filtering that efficiently handles dynamic keywords by focusing on the semantics of the words. We capture both semantic and textual features of the short-text messages based on 13 pre-defined intention labels. Moreover, the contextual embeddings of the texts are generated using various pre-trained NLP (Natural Language Processing) models. Finally, intention scores are computed for the pre-defined labels and a bunch of supervised learning classifiers are employed for filtering as spam or ham. Our approaches are evaluated on the SMS Spam Collection [ 24 ] benchmark dataset, and extensive experimentation shows interesting results. Our model did remarkably well with an accuracy of 98.07%, Precision and Recall of ∼ 0.97, which is better than few of the existing state-of-the-art alternatives. Though the accuracy of our approach is not the best among other existing approaches, the model is highly stable due to its emphasis on extracting the contextual features from the text through intention labels.

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Networks and Communications

Link

https://dl.acm.org/doi/pdf/10.1145/3538491

Reference33 articles.

1. A review of soft techniques for SMS spam classification: Methods, approaches and applications

2. A Review on Mobile SMS Spam Filtering Techniques

3. Mining Text Data

4. Contributions to the study of SMS spam filtering

5. A survey of learning-based techniques of email spam filtering

Cited by 20 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. An integrated model based on deep learning classifiers and pre-trained transformer for phishing URL detection;Future Generation Computer Systems;2024-12

2. INCEPT: A Framework for Duplicate Posts Classification with Combined Text Representations;ACM Transactions on the Web;2024-08-16

3. Multilingual SMS Spam Detection using BERT and LSTM;2024 International Conference on Innovations and Challenges in Emerging Technologies (ICICET);2024-06-07

4. From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language Models;2024 IEEE Symposium on Security and Privacy (SP);2024-05-19

5. On SMS Phishing Tactics and Infrastructure;2024 IEEE Symposium on Security and Privacy (SP);2024-05-19