Effective preprocessing based neural machine translation for English to Telugu cross-language information retrieval
-
Published:2021-06-01
Issue:2
Volume:10
Page:306
-
ISSN:2252-8938
-
Container-title:IAES International Journal of Artificial Intelligence (IJ-AI)
-
language:
-
Short-container-title:IJ-AI
Author:
Raju B. N. V. Narasimha,Raju M. S. V. S. Bhadri,V. Satyanarayana K. V.
Abstract
<span id="docs-internal-guid-5b69f940-7fff-f443-1f09-a00e5e983714"><span>In cross-language information retrieval (CLIR), the neural machine translation (NMT) plays a vital role. CLIR retrieves the information written in a language which is different from the user's query language. In CLIR, the main concern is to translate the user query from the source language to the target language. NMT is useful for translating the data from one language to another. NMT has better accuracy for different languages like English to German and so-on. In this paper, NMT has applied for translating English to Indian languages, especially for Telugu. Besides NMT, an effort is also made to improve accuracy by applying effective preprocessing mechanism. The role of effective preprocessing in improving accuracy will be less but countable. Machine translation (MT) is a data-driven approach where parallel corpus will act as input in MT. NMT requires a massive amount of parallel corpus for performing the translation. Building an English - Telugu parallel corpus is costly because they are resource-poor languages. Different mechanisms are available for preparing the parallel corpus. The major issue in preparing parallel corpus is data replication that is handled during preprocessing. The other issue in machine translation is the out-of-vocabulary (OOV) problem. Earlier dictionaries are used to handle OOV problems. To overcome this problem the rare words are segmented into sequences of subwords during preprocessing. The parameters like accuracy, perplexity, cross-entropy and BLEU scores shows better translation quality for NMT with effective preprocessing.</span></span>
Publisher
Institute of Advanced Engineering and Science
Subject
Electrical and Electronic Engineering,Artificial Intelligence,Information Systems and Management,Control and Systems Engineering
Cited by
3 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献
1. Translation English to Punjabi: A Concise Review of Significant Approaches;2024 International Conference on Computational Intelligence and Computing Applications (ICCICA);2024-05-23
2. Neural Machine Translation and Detailed Analysis of Impact of Pre-Processing Techniques;2023 3rd International Conference on Advance Computing and Innovative Technologies in Engineering (ICACITE);2023-05-12
3. KanSan: Kannada-Sanskrit Parallel Corpus Construction for Machine Translation;Communications in Computer and Information Science;2023