Identification of offensive language in Urdu using semantic and embedding models-Reference-Cited by-同舟云学术

Identification of offensive language in Urdu using semantic and embedding models

Published:2022-12-12 Issue: Volume:8 Page:e1169
ISSN:2376-5992
Container-title:PeerJ Computer Science
language:en
Short-container-title:

Author:

Hussain Sajid,Malik Muhammad Shahid Iqbal,Masood Nayyer

Abstract

Automatic identification of offensive/abusive language is very necessary to get rid of unwanted behavior. However, it is more challenging to generalize the solution due to the different grammatical structures and vocabulary of each language. Most of the prior work targeted western languages, however, one study targeted a low-resource language (Urdu). The prior study used basic linguistic features and a small dataset. This study designed a new dataset (collected from popular Pakistani Facebook pages) containing 7,500 posts for offensive language detection in Urdu. The proposed methodology used four types of feature engineering models: three are frequency-based and the fourth one is the embedding model. Frequency-based are either determined by the term frequency-inverse document frequency (TF-IDF) or bag-of-words or word n-gram feature vectors. The fourth is generated by the word2vec model, trained on the Urdu embeddings using a corpus of 196,226 Facebook posts. The experiments demonstrate that the stacking-based ensemble model with word2vec shows the best performance as a standalone model by achieving 88.27% accuracy. In addition, the wrapper-based feature selection method further improves performance. The hybrid combination of TF-IDF, bag-of-words, and word2vec feature models achieved 90% accuracy and 97% AUC. In addition, it outperformed the baseline with an improvement of 3.55% in accuracy, 3.68% in the recall, 3.60% in f1-measure, 3.67% in precision, and 2.71% in AUC. The findings of this research provide practical implications for commercial applications and future research.

Publisher

PeerJ

Subject

General Computer Science

Link

https://peerj.com/articles/cs-1169.pdf

Reference43 articles.

1. An information-theoretic perspective of tf–idf measures;Aizawa;Information Processing & Management,2003

2. Automatic detection of offensive language for urdu and roman urdu;Akhter;IEEE Access,2020

3. A new feature selection method for enhancing cancer diagnosis based on DNA microarray;Atlam,2020

4. Overview of the EVALITA 2018 hate speech detection task;Bosco,2018

5. A corpus of Turkish offensive language on social media;Çöltekin,2020

Cited by 12 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Detection of violence incitation expressions in Urdu tweets using convolutional neural network;Expert Systems with Applications;2024-07

2. Categorization of tweets for damages: infrastructure and human damage assessment using fine-tuned BERT model;PeerJ Computer Science;2024-02-16

3. Hate Speech and Target Community Detection in Nastaliq Urdu Using Transfer Learning Techniques;IEEE Access;2024

4. Effectiveness of ELMo embeddings, and semantic models in predicting review helpfulness;INTELL DATA ANAL;2024

5. Abusive Language Detection in Urdu Text: Leveraging Deep Learning and Attention Mechanism;IEEE Access;2024