Affiliation:
1. Kyoto University and Universitas Islam Riau, Riau, Indonesia
2. Kyoto University, Kyoto, Japan
Abstract
The lack or absence of parallel and comparable corpora makes bilingual lexicon extraction a difficult task for low-resource languages. The pivot language and cognate recognition approaches have been proven useful for inducing bilingual lexicons for such languages. We propose constraint-based bilingual lexicon induction for closely related languages by extending constraints from the recent pivot-based induction technique and further enabling multiple symmetry assumption cycle to reach many more cognates in the transgraph. We further identify cognate synonyms to obtain many-to-many translation pairs. This article utilizes four datasets: one Austronesian low-resource language and three Indo-European high-resource languages. We use three constraint-based methods from our previous work, the Inverse Consultation method and translation pairs generated from Cartesian product of input dictionaries as baselines. We evaluate our result using the metrics of precision, recall, and F-score. Our customizable approach allows the user to conduct cross validation to predict the optimal hyperparameters (cognate threshold and cognate synonym threshold) with various combination of heuristics and number of symmetry assumption cycles to gain the highest F-score. Our proposed methods have statistically significant improvement of precision and F-score compared to our previous constraint-based methods. The results show that our method demonstrates the potential to complement other bilingual dictionary creation methods like word alignment models using parallel corpora for high-resource languages while well handling low-resource languages.
Funder
Grant-in-Aid for Scientific Research
Indonesia Endownment Fund for Education
Grant-in-Aid for Young Scientists
Japan Society for the Promotion of Science
Publisher
Association for Computing Machinery (ACM)
Reference37 articles.
1. Carlos Ansótegui María Luisa Bonet and Jordi Levy. 2009. Solving (weighted) partial MaxSAT through satisfiability testing. In Theory and Applications of Satisfiability Testing-SAT 2009. Springer 427--440. Carlos Ansótegui María Luisa Bonet and Jordi Levy. 2009. Solving (weighted) partial MaxSAT through satisfiability testing. In Theory and Applications of Satisfiability Testing-SAT 2009. Springer 427--440.
2. Armin Biere Marijn Heule and Hans van Maaren. 2009. Handbook of Satisfiability. Vol. 185. IOS Press. Armin Biere Marijn Heule and Hans van Maaren. 2009. Handbook of Satisfiability. Vol. 185. IOS Press.
3. Lyle Campbell. 2013. Historical Linguistics. Edinburgh University Press. Lyle Campbell. 2013. Historical Linguistics. Edinburgh University Press.
4. Lyle Campbell and William J. Poser. 2008. Language classification. History and Method. Cambridge University Press Cambridge (2008). 10.1017/CBO9780511486906 Lyle Campbell and William J. Poser. 2008. Language classification. History and Method. Cambridge University Press Cambridge (2008). 10.1017/CBO9780511486906
Cited by
16 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献