Contextual Urdu Lemmatization Using Recurrent Neural Network Models-Reference-Cited by-同舟云学术

Contextual Urdu Lemmatization Using Recurrent Neural Network Models

Published:2023-01-13 Issue:2 Volume:11 Page:435
ISSN:2227-7390
Container-title:Mathematics
language:en
Short-container-title:Mathematics

Author:

Hafeez Rabab,Anwar Muhammad Waqas^ORCID,Jamal Muhammad Hasan^ORCID,Fatima Tayyaba,Espinosa Julio César Martínez,López Luis Alonso Dzul^ORCID,Thompson Ernesto Bautista,Ashraf Imran^ORCID

Abstract

In the field of natural language processing, machine translation is a colossally developing research area that helps humans communicate more effectively by bridging the linguistic gap. In machine translation, normalization and morphological analyses are the first and perhaps the most important modules for information retrieval (IR). To build a morphological analyzer, or to complete the normalization process, it is important to extract the correct root out of different words. Stemming and lemmatization are techniques commonly used to find the correct root words in a language. However, a few studies on IR systems for the Urdu language have shown that lemmatization is more effective than stemming due to infixes found in Urdu words. This paper presents a lemmatization algorithm based on recurrent neural network models for the Urdu language. However, lemmatization techniques for resource-scarce languages such as Urdu are not very common. The proposed model is trained and tested on two datasets, namely, the Urdu Monolingual Corpus (UMC) and the Universal Dependencies Corpus of Urdu (UDU). The datasets are lemmatized with the help of recurrent neural network models. The Word2Vec model and edit trees are used to generate semantic and syntactic embedding. Bidirectional long short-term memory (BiLSTM), bidirectional gated recurrent unit (BiGRU), bidirectional gated recurrent neural network (BiGRNN), and attention-free encoder–decoder (AFED) models are trained under defined hyperparameters. Experimental results show that the attention-free encoder-decoder model achieves an accuracy, precision, recall, and F-score of 0.96, 0.95, 0.95, and 0.95, respectively, and outperforms existing models.

Funder

European University of Atlantic

Publisher

MDPI AG

Subject

General Mathematics,Engineering (miscellaneous),Computer Science (miscellaneous)

Link

https://www.mdpi.com/2227-7390/11/2/435/pdf

Reference29 articles.

1. Method of lemmatizer selections in multiplexing lemmatization;Sychev;IOP Conf. Ser. Mater. Sci. Eng.,2019

2. A hybrid approach for Arabic lemmatization;Boudchiche;Int. J. Speech Technol.,2019

3. Samir, A., and Lahbib, Z. (2018, January 4–5). Stemming and lemmatization for information retrieval systems in amazigh language. Proceedings of the International Conference on Big Data, Cloud and Applications, Kenitra, Morocco.

4. STEMUR: An Automated Word Conflation Algorithm for the Urdu Language;Fatima;Trans. Asian Low-Resour. Lang. Inf. Process.,2021

5. A survey on Urdu and Urdu like language stemmers and stemming techniques;Jabbar;Artif. Intell. Rev.,2018

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A Systematic Review of Computational Approaches to Deciphering Bronze Age Aegean and Cypriot Scripts;Computational Linguistics;2024

2. Developing an Urdu Lemmatizer Using a Dictionary-Based Lookup Approach;Applied Sciences;2023-04-19

3. Modeling Topics in DFA-Based Lemmatized Gujarati Text;Sensors;2023-03-01