A Bayesian Alignment Approach to Transliteration Mining-Reference-Cited by-同舟云学术

A Bayesian Alignment Approach to Transliteration Mining

Published:2013-08 Issue:3 Volume:12 Page:1-22
ISSN:1530-0226
Container-title:ACM Transactions on Asian Language Information Processing
language:en
Short-container-title:ACM Transactions on Asian Language Information Processing

Author:

Fukunishi Takaaki¹,Finch Andrew²,Yamamoto Seiichi¹,Sumita Eiichiro²

Affiliation:

1. Doshisha University

2. NICT

Abstract

In this article we present a technique for mining transliteration pairs using a set of simple features derived from a many-to-many bilingual forced-alignment at the grapheme level to classify candidate transliteration word pairs as correct transliterations or not. We use a nonparametric Bayesian method for the alignment process, as this process rewards the reuse of parameters, resulting in compact models that align in a consistent manner and tend not to over-fit. Our approach uses the generative model resulting from aligning the training data to force-align the test data. We rely on the simple assumption that correct transliteration pairs would be well modeled and generated easily, whereas incorrect pairs---being more random in character---would be more costly to model and generate. Our generative model generates by concatenating bilingual grapheme sequence pairs. The many-to-many generation process is essential for handling many languages with non-Roman scripts, and it is hard to train well using a maximum likelihood techniques, as these tend to over-fit the data. Our approach works on the principle that generation using only grapheme sequence pairs that are in the model results in a high probability derivation, whereas if the model is forced to introduce a new parameter in order to explain part of the candidate pair, the derivation probability is substantially reduced and severely reduced if the new parameter corresponds to a sequence pair composed of a large number of graphemes. The features we extract from the alignment of the test data are not only based on the scores from the generative model, but also on the relative proportions of each sequence that are hard to generate. The features are used in conjunction with a support vector machine classifier trained on known positive examples together with synthetic negative examples to determine whether a candidate word pair is a correct transliteration pair. In our experiments, we used all data tracks from the 2010 Named-Entity Workshop (NEWS’10) and use the performance of the best system for each language pair as a reference point. Our results show that the new features we propose are powerfully predictive, enabling our approach to achieve levels of performance on this task that are comparable to the state of the art.

Publisher

Association for Computing Machinery (ACM)

Subject

General Computer Science

Link

https://dl.acm.org/doi/pdf/10.1145/2499955.2499957

Reference34 articles.

1. An improved error model for noisy channel spelling correction

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Agreement on Target-Bidirectional Recurrent Neural Networks for Sequence-to-Sequence Learning;Journal of Artificial Intelligence Research;2020-03-19

2. Segmentation and Alignment of Chinese and Khmer Bilingual Names Based on Hierarchical Dirichlet Process;Advances in Intelligent Systems and Computing;2018-10-05

3. Machine transliteration and transliterated text retrieval: a survey;Sādhanā;2018-06

4. Inducing a Bilingual Lexicon from Short Parallel Multiword Sequences;ACM Transactions on Asian and Low-Resource Language Information Processing;2017-04-06

5. Noise-aware Character Alignment for Extracting Transliteration Fragments;Journal of Natural Language Processing;2014