Pseudo Relevance Feedback with Deep Language Models and Dense Retrievers: Successes and Pitfalls-Reference-Cited by-同舟云学术

Pseudo Relevance Feedback with Deep Language Models and Dense Retrievers: Successes and Pitfalls

Published:2023-04-10 Issue:3 Volume:41 Page:1-40
ISSN:1046-8188
Container-title:ACM Transactions on Information Systems
language:en
Short-container-title:ACM Trans. Inf. Syst.

Author:

Li Hang¹^ORCID,Mourad Ahmed¹^ORCID,Zhuang Shengyao¹^ORCID,Koopman Bevan²^ORCID,Zuccon Guido¹^ORCID

Affiliation:

1. IElab, The University of Queensland, Queensland, Australia

2. Australian E-Health Research Centre, CSIRO, Queensland, Australia

Abstract

Pseudo Relevance Feedback (PRF) is known to improve the effectiveness of bag-of-words retrievers. At the same time, deep language models have been shown to outperform traditional bag-of-words rerankers. However, it is unclear how to integrate PRF directly with emergent deep language models. This article addresses this gap by investigating methods for integrating PRF signals with rerankers and dense retrievers based on deep language models. We consider text-based, vector-based and hybrid PRF approaches and investigate different ways of combining and scoring relevance signals. An extensive empirical evaluation was conducted across four different datasets and two task settings (retrieval and ranking). Text-based PRF results show that the use of PRF had a mixed effect on deep rerankers across different datasets. We found that the best effectiveness was achieved when (i) directly concatenating each PRF passage with the query, searching with the new set of queries, and then aggregating the scores; (ii) using Borda to aggregate scores from PRF runs. Vector-based PRF results show that the use of PRF enhanced the effectiveness of deep rerankers and dense retrievers over several evaluation metrics. We found that higher effectiveness was achieved when (i) the query retains either the majority or the same weight within the PRF mechanism, and (ii) a shallower PRF signal (i.e., a smaller number of top-ranked passages) was employed, rather than a deeper signal. Our vector-based PRF method is computationally efficient; thus, this represents a general PRF method others can use with deep rerankers and dense retrievers.

Funder

Grain Research and Development Corporation project AgAsk

Australian Research Council DECRA Research Fellowship

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Science Applications,General Business, Management and Accounting,Information Systems

Link

https://dl.acm.org/doi/pdf/10.1145/3570724

Reference83 articles.

1. UMass at TREC 2004: Novelty and HARD;Abdul-Jaleel Nasreen;Computer Science Department Faculty Publication Series,2004

2. Predicting Efficiency/Effectiveness Trade-offs for Dense vs. Sparse Retrieval Strategy Selection

3. Models for metasearch

4. Query expansion techniques for information retrieval: A survey

5. Selecting good expansion terms for pseudo-relevance feedback

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Learning to Jointly Transform and Rank Difficult Queries;Lecture Notes in Computer Science;2024

2. GenQREnsemble: Zero-Shot LLM Ensemble Prompting for Generative Query Reformulation;Lecture Notes in Computer Science;2024

3. Augmenting Passage Representations with Query Generation for Enhanced Cross-Lingual Dense Retrieval;Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval;2023-07-18