Exhaustive Study into Machine Learning and Deep Learning Methods for Multilingual Cyberbullying Detection in Bangla and Chittagonian Texts-Reference-Cited by-同舟云学术

Exhaustive Study into Machine Learning and Deep Learning Methods for Multilingual Cyberbullying Detection in Bangla and Chittagonian Texts

Published:2024-04-26 Issue:9 Volume:13 Page:1677
ISSN:2079-9292
Container-title:Electronics
language:en
Short-container-title:Electronics

Author:

Mahmud Tanjim¹²^ORCID,Ptaszynski Michal¹^ORCID,Masui Fumito¹^ORCID

Affiliation:

1. Text Information Processing Laboratory, Kitami Institute of Technology, 165 Koen-cho, Kitami City 090-8507, Hokkaido, Japan

2. Department of Computer Science and Engineering, Rangamati Science and Technology University, Rangamati 4500, Bangladesh

Abstract

Cyberbullying is a serious problem in online communication. It is important to find effective ways to detect cyberbullying content to make online environments safer. In this paper, we investigated the identification of cyberbullying contents from the Bangla and Chittagonian languages, which are both low-resource languages, with the latter being an extremely low-resource language. In the study, we used both traditional baseline machine learning methods, as well as a wide suite of deep learning methods especially focusing on hybrid networks and transformer-based multilingual models. For the data, we collected over 5000 both Bangla and Chittagonian text samples from social media. Krippendorff’s alpha and Cohen’s kappa were used to measure the reliability of the dataset annotations. Traditional machine learning methods used in this research achieved accuracies ranging from 0.63 to 0.711, with SVM emerging as the top performer. Furthermore, employing ensemble models such as Bagging with 0.70 accuracy, Boosting with 0.69 accuracy, and Voting with 0.72 accuracy yielded promising results. In contrast, deep learning models, notably CNN, achieved accuracies ranging from 0.69 to 0.811, thus outperforming traditional ML approaches, with CNN exhibiting the highest accuracy. We also proposed a series of hybrid network-based models, including BiLSTM+GRU with an accuracy of 0.799, CNN+LSTM with 0.801 accuracy, CNN+BiLSTM with 0.78 accuracy, and CNN+GRU with 0.804 accuracy. Notably, the most complex model, (CNN+LSTM)+BiLSTM, attained an accuracy of 0.82, thus showcasing the efficacy of hybrid architectures. Furthermore, we explored transformer-based models, such as XLM-Roberta with 0.841 accuracy, Bangla BERT with 0.822 accuracy, Multilingual BERT with 0.821 accuracy, BERT with 0.82 accuracy, and Bangla ELECTRA with 0.785 accuracy, which showed significantly enhanced accuracy levels. Our analysis demonstrates that deep learning methods can be highly effective in addressing the pervasive issue of cyberbullying in several different linguistic contexts. We show that transformer models can efficiently circumvent the language dependence problem that plagues conventional transfer learning methods. Our findings suggest that hybrid approaches and transformer-based embeddings can effectively tackle the problem of cyberbullying across online platforms.

Publisher

MDPI AG

Link

https://www.mdpi.com/2079-9292/13/9/1677/pdf

Reference86 articles.

1. (2023, January 15). Bangladesh Telecommunication Regulatory Commission, Available online: http://www.btrc.gov.bd/site/page/347df7fe-409f-451e-a415-65b109a207f5/-.

2. (2023, January 20). United Nations Development Programme. Available online: https://www.undp.org/bangladesh/blog/digital-bangladesh-innovative-bangladesh-road-2041.

3. (2023, April 01). Chittagong City in Bangladesh. Available online: https://en.wikipedia.org/wiki/Chittagong.

4. (2023, April 28). StatCounter Global Stats. Available online: https://gs.statcounter.com/social-media-stats/all/bangladesh/#monthly-202203-202303.

5. (2023, February 11). Chittagonian Language. Available online: https://en.wikipedia.org/wiki/Chittagonian_language.

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Bias and Cyberbullying Detection and Data Generation Using Transformer Artificial Intelligence Models and Top Large Language Models;Electronics;2024-08-29

2. Investigating the Effectiveness of Deep Learning and Machine Learning for Bangla Poems Genre Classification;2023 4th International Conference on Intelligent Technologies (CONIT);2024-06-21

3. Hybrid Deep Transfer Learning Framework for Humerus Fracture Detection and Classification from X-ray Images;2023 4th International Conference on Intelligent Technologies (CONIT);2024-06-21

4. Enhancing Cyberbullying Detection on Social Media Using Transformer Models;2024 5th Technology Innovation Management and Engineering Science International Conference (TIMES-iCON);2024-06-19

5. Analyzing Sentiments in eLearning: A Comparative Study of Bangla and Romanized Bangla Text Using Transformers;IEEE Access;2024