A Comprehensive Analysis of Bilingual Lexicon Induction-Reference-Cited by-同舟云学术

A Comprehensive Analysis of Bilingual Lexicon Induction

Published:2017-06 Issue:2 Volume:43 Page:273-310
ISSN:0891-2017
Container-title:Computational Linguistics
language:en
Short-container-title:Computational Linguistics

Author:

Irvine Ann¹,Callison-Burch Chris²

Affiliation:

1. Johns Hopkins University

2. University of Pennsylvania

Abstract

Bilingual lexicon induction is the task of inducing word translations from monolingual corpora in two languages. In this article we present the most comprehensive analysis of bilingual lexicon induction to date. We present experiments on a wide range of languages and data sizes. We examine translation into English from 25 foreign languages: Albanian, Azeri, Bengali, Bosnian, Bulgarian, Cebuano, Gujarati, Hindi, Hungarian, Indonesian, Latvian, Nepali, Romanian, Serbian, Slovak, Somali, Spanish, Swedish, Tamil, Telugu, Turkish, Ukrainian, Uzbek, Vietnamese, and Welsh. We analyze the behavior of bilingual lexicon induction on low-frequency words, rather than testing solely on high-frequency words, as previous research has done. Low-frequency words are more relevant to statistical machine translation, where systems typically lack translations of rare words that fall outside of their training data. We systematically explore a wide range of features and phenomena that affect the quality of the translations discovered by bilingual lexicon induction. We provide illustrative examples of the highest ranking translations for orthogonal signals of translation equivalence like contextual similarity and temporal similarity. We analyze the effects of frequency and burstiness, and the sizes of the seed bilingual dictionaries and the monolingual training corpora. Additionally, we introduce a novel discriminative approach to bilingual lexicon induction. Our discriminative model is capable of combining a wide variety of features that individually provide only weak indications of translation equivalence. When feature weights are discriminatively set, these signals produce dramatically higher translation quality than previous approaches that combined signals in an unsupervised fashion (e.g., using minimum reciprocal rank). We also directly compare our model's performance against a sophisticated generative approach, the matching canonical correlation analysis (MCCA) algorithm used by Haghighi et al. ( 2008 ). Our algorithm achieves an accuracy of 42% versus MCCA's 15%.

Publisher

MIT Press - Journals

Subject

Artificial Intelligence,Computer Science Applications,Linguistics and Language,Language and Linguistics

Link

https://www.mitpressjournals.org/doi/pdf/10.1162/COLI_a_00284

Reference50 articles.

1. Exploiting comparable corpora with TER and TERp

2. Abdul-Rauf, Sadaf and Holger Schwenk. 2009b. On the use of comparable corpora to improve SMT performance. In Proceedings of the 12th Conference of the European Chapter of the ACL (EACL 2009), pages 16–23, Athens.

3. Agarwal, Alekh, Oliveier Chapelle, Miroslav Dudík, and John Langford. 2014. A reliable effective terascale linear learning system. Journal of Machine Learning Research, 15:1111–1133.

4. Gazpacho and summer rash

5. Statistical Extraction and Comparison of Pivot Words for Bilingual Lexicon Extension

Cited by 12 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Multilingual question answering systems for knowledge graphs – a survey;Semantic Web;2024-08-28

2. The hypergeometric test performs comparably to TF-IDF on standard text analysis tasks;Multimedia Tools and Applications;2023-09-08

3. Other Applications of Comparable Corpora;Building and Using Comparable Corpora for Multilingual Natural Language Processing;2023

4. Induction of Bilingual Dictionaries;Building and Using Comparable Corpora for Multilingual Natural Language Processing;2023

5. Extraction of Parallel Sentences;Building and Using Comparable Corpora for Multilingual Natural Language Processing;2023