A Lightweight Multi-View Learning Approach for Phishing Attack Detection Using Transformer with Mixture of Experts-Reference-Cited by-同舟云学术

A Lightweight Multi-View Learning Approach for Phishing Attack Detection Using Transformer with Mixture of Experts

Published:2023-06-22 Issue:13 Volume:13 Page:7429
ISSN:2076-3417
Container-title:Applied Sciences
language:en
Short-container-title:Applied Sciences

Author:

Wang Yanbin¹,Ma Wenrui¹,Xu Haitao¹,Liu Yiwei²,Yin Peng²³

Affiliation:

1. School of Cyber and Technology, Zhejiang University, Hangzhou 310027, China

2. Defence Industry Secrecy Examination and Certification Center, Beijing 100089, China

3. School of Cyber Security, University of Chinese Academy of Sciences, Beijing 101408, China

Abstract

Phishing poses a significant threat to the financial and privacy security of internet users and often serves as the starting point for cyberattacks. Many machine-learning-based methods for detecting phishing websites rely on URL analysis, offering simplicity and efficiency. However, these approaches are not always effective due to the following reasons: (1) highly concealed phishing websites may employ tactics such as masquerading URL addresses to deceive machine learning models, and (2) phishing attackers frequently change their phishing website URLs to evade detection. In this study, we propose a robust, multi-view Transformer model with an expert-mixture mechanism for accurate phishing website detection utilizing website URLs, attributes, content, and behavioral information. Specifically, we first adapted a pretrained language model for URL representation learning by applying adversarial post-training learning in order to extract semantic information from URLs. Next, we captured the attribute, content, and behavioral features of the websites and encoded them as vectors, which, alongside the URL embeddings, constitute the website’s multi-view information. Subsequently, we introduced a mixture-of-experts mechanism into the Transformer network to learn knowledge from different views and adaptively fuse information from various views. The proposed method outperforms state-of-the-art approaches in evaluations of real phishing websites, demonstrating greater performance with less label dependency. Furthermore, we show the superior robustness and enhanced adaptability of the proposed method to unseen samples and data drift in more challenging experimental settings.

Publisher

MDPI AG

Subject

Fluid Flow and Transfer Processes,Computer Science Applications,Process Chemistry and Technology,General Engineering,Instrumentation,General Materials Science

Link

https://www.mdpi.com/2076-3417/13/13/7429/pdf

Reference49 articles.

1. Zabihimayvan, M., and Doran, D. (2019, January 23–26). Fuzzy rough set feature selection to enhance phishing attack detection. Proceedings of the 2019 IEEE International Conference on Fuzzy Systems (FUZZ-IEEE), New Orleans, LA, USA.

2. Detection of phishing attacks: A machine learning approach;Basnet;Soft Comput. Appl. Ind.,2008

3. A deep learning technique for web phishing detection combined URL features and visual similarity;Int. J. Comput. Netw. Commun. (IJCNC),2020

4. Cui, Q., Jourdan, G.V., Bochmann, G.V., Couturier, R., and Onut, I.V. (2017, January 3–7). Tracking phishing attacks over time. Proceedings of the 26th International Conference on World Wide Web, Perth, Australia.

5. Mobile phishing attacks and defence mechanisms: State of art and open research challenges;Goel;Comput. Secur.,2018

Cited by 10 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. PMANet: Malicious URL detection via post-trained language model guided multi-level feature attention network;Information Fusion;2025-01

2. TransURL: Improving malicious URL detection with multi-layer Transformer encoding and multi-scale pyramid features;Computer Networks;2024-11

3. Improved Generative Adversarial Network for Phishing Attack Detection;2024 4th International Conference on Pervasive Computing and Social Networking (ICPCSN);2024-05-03

4. Enhancing Network Attack Detection Accuracy through the Integration of Large Language Models and Synchronized Attention Mechanism;Applied Sciences;2024-04-30

5. An Improved Webpage Phishing Detection Based on an Ensemble Feature Selection and Multilevel Techniques;2024 International Conference on Science, Engineering and Business for Driving Sustainable Development Goals (SEB4SDG);2024-04-02