Framework for Handling Rare Word Problems in Neural Machine Translation System Using Multi-Word Expressions-Reference-Cited by-同舟云学术

Framework for Handling Rare Word Problems in Neural Machine Translation System Using Multi-Word Expressions

Published:2022-10-31 Issue:21 Volume:12 Page:11038
ISSN:2076-3417
Container-title:Applied Sciences
language:en
Short-container-title:Applied Sciences

Author:

Garg Kamal Deep^ORCID,Shekhar Shashi,Kumar Ajit,Goyal Vishal,Sharma Bhisham^ORCID,Chengoden Rajeswari^ORCID,Srivastava Gautam^ORCID

Abstract

Machine Translation (MT) systems are now being improved with the use of an ongoing methodology known as Neural Machine Translation (NMT). Natural language processing (NLP) researchers have shown that NMT systems are unable to deal with out-of-vocabulary (OOV) words and multi-word expressions (MWEs) in the text. OOV terms are those that are not currently included in the vocabulary that is used by the NMT system. MWEs are phrases that consist of a minimum of two terms but are treated as a single unit. MWEs have great importance in NLP, linguistic theory, and MT systems. In this article, OOV words and MWEs are handled for the Punjabi to English NMT system. A parallel corpus for Punjabi to English containing MWEs was developed and used to train the different models of NMT. Punjabi is a low-resource language as it lacks the availability of a large parallel corpus for building various NLP tools, and this is an attempt to improve the accuracy of Punjabi in the English NMT system by using named entities and MWEs in the corpus. The developed NMT models were assessed using human evaluation through adequacy, fluency and overall rating as well as automated assessment tools such as the bilingual evaluation study (BLEU) and translation error rate (TER) score. Results show that using word embedding (WE) and MWEs corpus increased the accuracy of translation for the Punjabi to English language pair. The best BLEU score obtained was 15.45 for the small test set, 43.32 for the medium test set, and 34.5 for the large test set, respectively. The best TER rate score obtained was 57.34% for the small test set, 37.29% for the medium test set, and 53.79% for the large test set, repectively.

Publisher

MDPI AG

Subject

Fluid Flow and Transfer Processes,Computer Science Applications,Process Chemistry and Technology,General Engineering,Instrumentation,General Materials Science

Link

https://www.mdpi.com/2076-3417/12/21/11038/pdf

Reference41 articles.

1. Hutchins, W.J. Machine Translation: A Brief History, 1995.

2. Review Article: Example-Based Machine Translation;Somers;Mach. Transl.,1999

3. Recurrent Continuous Translation Models. EMNLP 2013–2013 Conference on Empirical Methods in Natural Language Processing;Kalchbrenner;Proc. Conf.,2013

4. Bone Cancer Detection Using Feature Extraction Based Machine Learning Model;Sharma;Comput. Math. Methods Med.,2021

5. Lahoura, V., Singh, H., Aggarwal, A., Sharma, B., Mohammed, M.A., Damaševičius, R., Kadry, S., and Cengiz, K. Cloud Computing-Based Framework for Breast Cancer Diagnosis Using Extreme Learning Machine. Diagnostics, 2021. 11.

Cited by 10 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Optimized BERT: an effective attention layer based deep learning technique utilizing for multiword term extraction;International Journal of Information Technology;2024-04-25

2. Translation Systems: A Synoptic Survey of Deep Learning Approaches and Techniques;2024 11th International Conference on Reliability, Infocom Technologies and Optimization (Trends and Future Directions) (ICRITO);2024-03-14

3. Ensuring Security of Data Through Transformation Based Encryption Algorithm in Image Steganography;Lecture Notes in Electrical Engineering;2024

4. Effective Spam Detection with Machine Learning;Croatian Regional Development Journal;2023-12-01

5. Analysis on the Legal System of International Technology Trade Management Based on Data Mining Analysis;International Journal on Semantic Web and Information Systems;2023-08-18