Text Classification Based on Convolutional Neural Networks and Word Embedding for Low-Resource Languages: Tigrinya-Reference-Cited by-同舟云学术

Text Classification Based on Convolutional Neural Networks and Word Embedding for Low-Resource Languages: Tigrinya

Published:2021-01-25 Issue:2 Volume:12 Page:52
ISSN:2078-2489
Container-title:Information
language:en
Short-container-title:Information

Author:

Fesseha Awet^ORCID,Xiong Shengwu,Emiru Eshete Derb^ORCID,Diallo Moussa^ORCID,Dahou Abdelghani

Abstract

This article studies convolutional neural networks for Tigrinya (also referred to as Tigrigna), which is a family of Semitic languages spoken in Eritrea and northern Ethiopia. Tigrinya is a “low-resource” language and is notable in terms of the absence of comprehensive and free data. Furthermore, it is characterized as one of the most semantically and syntactically complex languages in the world, similar to other Semitic languages. To the best of our knowledge, no previous research has been conducted on the state-of-the-art embedding technique that is shown here. We investigate which word representation methods perform better in terms of learning for single-label text classification problems, which are common when dealing with morphologically rich and complex languages. Manually annotated datasets are used here, where one contains 30,000 Tigrinya news texts from various sources with six categories of “sport”, “agriculture”, “politics”, “religion”, “education”, and “health” and one unannotated corpus that contains more than six million words. In this paper, we explore pretrained word embedding architectures using various convolutional neural networks (CNNs) to predict class labels. We construct a CNN with a continuous bag-of-words (CBOW) method, a CNN with a skip-gram method, and CNNs with and without word2vec and FastText to evaluate Tigrinya news articles. We also compare the CNN results with traditional machine learning models and evaluate the results in terms of the accuracy, precision, recall, and F1 scoring techniques. The CBOW CNN with word2vec achieves the best accuracy with 93.41%, significantly improving the accuracy for Tigrinya news classification.

Publisher

MDPI AG

Subject

Information Systems

Link

https://www.mdpi.com/2078-2489/12/2/52/pdf

Reference47 articles.

1. A comprehensive survey of arabic sentiment analysis

2. The Origin and Development of Tigrinya Language Publications (1886-1991) Volume Onehttps://scholarcommons.scu.edu/cgi/viewcontent.cgi?article=1130&context=library

3. Tigrinya Part-of-Speech Tagging with Morphological Patterns and the New Nagaoka Tigrinya Corpus

4. Tigrinya Morphological Segmentation with Bidirectional Long Short-Term Memory Neural Networks and its Effect on English-Tigrinya Machine Translation;Tedla,2018

Cited by 63 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. The analysis of art design under improved convolutional neural network based on the Internet of Things technology;Scientific Reports;2024-09-10

2. Automatic Extraction and Cluster Analysis of Natural Disaster Metadata Based on the Unified Metadata Framework;ISPRS International Journal of Geo-Information;2024-06-14

3. Speaker-based language identification for Ethio-Semitic languages using CRNN and hybrid features;Network: Computation in Neural Systems;2024-06-04

4. Advanced Ensemble Classifier Techniques for Predicting Tumor Viability in Osteosarcoma Histological Slide Images;Applied Data Science and Analysis;2024-05-29

5. Text Representation Based on WT-GloVe Word Vector Weighting Model;2024 International Conference on Intelligent Computing and Robotics (ICICR);2024-04-12