Multilingual Hate Speech Detection: A Semi-Supervised Generative Adversarial Approach-Reference-Cited by-同舟云学术

Multilingual Hate Speech Detection: A Semi-Supervised Generative Adversarial Approach

Published:2024-04-18 Issue:4 Volume:26 Page:344
ISSN:1099-4300
Container-title:Entropy
language:en
Short-container-title:Entropy

Author:

Mnassri Khouloud¹^ORCID,Farahbakhsh Reza¹^ORCID,Crespi Noel¹

Affiliation:

1. Samovar, Télécom SudParis, Institut Polytechnique de Paris, 91120 Palaiseau, France

Abstract

Social media platforms have surpassed cultural and linguistic boundaries, thus enabling online communication worldwide. However, the expanded use of various languages has intensified the challenge of online detection of hate speech content. Despite the release of multiple Natural Language Processing (NLP) solutions implementing cutting-edge machine learning techniques, the scarcity of data, especially labeled data, remains a considerable obstacle, which further requires the use of semisupervised approaches along with Generative Artificial Intelligence (Generative AI) techniques. This paper introduces an innovative approach, a multilingual semisupervised model combining Generative Adversarial Networks (GANs) and Pretrained Language Models (PLMs), more precisely mBERT and XLM-RoBERTa. Our approach proves its effectiveness in the detection of hate speech and offensive language in Indo-European languages (in English, German, and Hindi) when employing only 20% annotated data from the HASOC2019 dataset, thereby presenting significantly high performances in each of multilingual, zero-shot crosslingual, and monolingual training scenarios. Our study provides a robust mBERT-based semisupervised GAN model (SS-GAN-mBERT) that outperformed the XLM-RoBERTa-based model (SS-GAN-XLM) and reached an average F1 score boost of 9.23% and an accuracy increase of 5.75% over the baseline semisupervised mBERT model.

Publisher

MDPI AG

Link

https://www.mdpi.com/1099-4300/26/4/344/pdf

Reference50 articles.

1. Language models are few-shot learners;Larochelle;Proceedings of the Advances in Neural Information Processing Systems,2020

2. Li, J., Tang, T., Zhao, W.X., Nie, J.Y., and Wen, J.R. (2022). Pretrained Language Models for Text Generation: A Survey. arXiv.

3. An Empirical Survey of Data Augmentation for Limited Data Learning in NLP;Chen;Trans. Assoc. Comput. Linguist.,2023

4. Generative ai;Feuerriegel;Bus. Inf. Syst. Eng.,2024

5. Multilingual use of Twitter: Social networks at the language frontier;Eleta;Comput. Hum. Behav.,2014

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Adversarial attacks and defenses for large language models (LLMs): methods, frameworks & challenges;International Journal of Multimedia Information Retrieval;2024-06-25