Using Sub-character Level Information for Neural Machine Translation of Logographic Languages-Reference-Cited by-同舟云学术

Using Sub-character Level Information for Neural Machine Translation of Logographic Languages

Published:2021-04-08 Issue:2 Volume:20 Page:1-15
ISSN:2375-4699
Container-title:ACM Transactions on Asian and Low-Resource Language Information Processing
language:en
Short-container-title:ACM Trans. Asian Low-Resour. Lang. Inf. Process.

Author:

Zhang Longtu¹^ORCID,Komachi Mamoru¹

Affiliation:

1. Tokyo Metropolitan University, Hino, Tokyo, Japan

Abstract

Logographic and alphabetic languages (e.g., Chinese vs. English) have different writing systems linguistically. Languages belonging to the same writing system usually exhibit more sharing information, which can be used to facilitate natural language processing tasks such as neural machine translation (NMT). This article takes advantage of the logographic characters in Chinese and Japanese by decomposing them into smaller units, thus more optimally utilizing the information these characters share in the training of NMT systems in both encoding and decoding processes. Experiments show that the proposed method can robustly improve the NMT performance of both “logographic” language pairs (JA–ZH) and “logographic + alphabetic” (JA–EN and ZH–EN) language pairs in both supervised and unsupervised NMT scenarios. Moreover, as the decomposed sequences are usually very long, extra position features for the transformer encoder can help with the modeling of these long sequences. The results also indicate that, theoretically, linguistic features can be manipulated to obtain higher share token rates and further improve the performance of natural language processing systems.

Publisher

Association for Computing Machinery (ACM)

Subject

General Computer Science

Link

https://dl.acm.org/doi/pdf/10.1145/3431727

Reference33 articles.

1. Steven Bird Ewan Klein and Edward Loper. 2009. Natural Language Processing with Python. O’Reilly. Steven Bird Ewan Klein and Edward Loper. 2009. Natural Language Processing with Python. O’Reilly.

Cited by 10 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Neural Machine Translation for Low-Resource Languages from a Chinese-centric Perspective: A Survey;ACM Transactions on Asian and Low-Resource Language Information Processing;2024-06-21

2. Sense-Aware Decoder for Character Based Japanese-Chinese NMT;IEICE Transactions on Information and Systems;2024-04-01

3. Sustained activation of NF-κB through constitutively active IKKβ leads to senescence bypass in murine dermal fibroblasts;Cell Cycle;2024-02

4. Variable Window and Deadline-Aware Sensor Attack Detector for Automotive CPS;2023 IEEE 26th International Symposium on Real-Time Distributed Computing (ISORC);2023-05

5. Cellular Senescence and Frailty in Transplantation;Current Transplantation Reports;2023-03-21