On cross-lingual retrieval with multilingual text encoders-Reference-Cited by-同舟云学术

On cross-lingual retrieval with multilingual text encoders

Published:2022-03-07 Issue:2 Volume:25 Page:149-183
ISSN:1386-4564
Container-title:Information Retrieval Journal
language:en
Short-container-title:Inf Retrieval J

Author:

Litschko Robert,Vulić Ivan,Ponzetto Simone Paolo,Glavaš Goran

Abstract

AbstractPretrained multilingual text encoders based on neural transformer architectures, such as multilingual BERT (mBERT) and XLM, have recently become a default paradigm for cross-lingual transfer of natural language processing models, rendering cross-lingual word embedding spaces (CLWEs) effectively obsolete. In this work we present a systematic empirical study focused on the suitability of the state-of-the-art multilingual encoders for cross-lingual document and sentence retrieval tasks across a number of diverse language pairs. We first treat these models as multilingual text encoders and benchmark their performance in unsupervised ad-hoc sentence- and document-level CLIR. In contrast to supervised language understanding, our results indicate that for unsupervised document-level CLIR—a setup with no relevance judgments for IR-specific fine-tuning—pretrained multilingual encoders on average fail to significantly outperform earlier models based on CLWEs. For sentence-level retrieval, we do obtain state-of-the-art performance: the peak scores, however, are met by multilingual encoders that have been further specialized, in a supervised fashion, for sentence understanding tasks, rather than using their vanilla ‘off-the-shelf’ variants. Following these results, we introduce localized relevance matching for document-level CLIR, where we independently score a query against document sections. In the second part, we evaluate multilingual encoders fine-tuned in a supervised fashion (i.e., we learn to rank) on English relevance data in a series of zero-shot language and domain transfer CLIR experiments. Our results show that, despite the supervision, and due to the domain and language shift, supervised re-ranking rarely improves the performance of multilingual transformers as unsupervised base rankers. Finally, only with in-domain contrastive fine-tuning (i.e., same domain, only language transfer), we manage to improve the ranking quality. We uncover substantial empirical differences between cross-lingual retrieval results and results of (zero-shot) cross-lingual transfer for monolingual retrieval in target languages, which point to “monolingual overfitting” of retrieval models trained on monolingual (English) data, even if they are based on multilingual transformers.

Funder

European Research Council

Ministerium für Wirtschaft, Arbeit und Wohnungsbau Baden-Württemberg

Universität Mannheim

Publisher

Springer Science and Business Media LLC

Subject

Library and Information Sciences,Information Systems

Link

https://link.springer.com/content/pdf/10.1007/s10791-022-09406-x.pdf

Reference86 articles.

1. Akkalyoncu Yilmaz, Z., Yang, W., Zhang, H., & Lin, J. (2019). Cross-domain modeling of sentence-level evidence for document retrieval. In Proceedings of EMNLP, pp. 3490–3496.

2. Artetxe, M., Labaka, G., & Agirre, E. (2018). A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings. In Proceedings of ACL, pp. 789–798.

3. Artetxe, M., & Schwenk, H. (2019). Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond. Transactions of the ACL pp. 597–610.

4. Beltagy, I., Peters, M.E., & Cohan, A. (2020). Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150

5. Braschler, M. (2003). CLEF 2003–Overview of results. In Workshop of the cross-language evaluation forum for european languages, pp. 44–63.

Cited by 9 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Steering Large Language Models for Cross-lingual Information Retrieval;Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval;2024-07-10

2. Multilingual Meta-Distillation Alignment for Semantic Retrieval;Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval;2024-07-10

3. Query in Your Tongue: Reinforce Large Language Models with Retrievers for Cross-lingual Search Generative Experience;Proceedings of the ACM Web Conference 2024;2024-05-13

4. Unsupervised multilingual machine translation with pretrained cross-lingual encoders;Knowledge-Based Systems;2024-01

5. Geographic Adaptation of Pretrained Language Models;Transactions of the Association for Computational Linguistics;2024