Boosting Search Performance Using Query Variations-Reference-Cited by-同舟云学术

Boosting Search Performance Using Query Variations

Published:2019-10-31 Issue:4 Volume:37 Page:1-25
ISSN:1046-8188
Container-title:ACM Transactions on Information Systems
language:en
Short-container-title:ACM Trans. Inf. Syst.

Author:

Benham Rodger¹^ORCID,Mackenzie Joel¹^ORCID,Moffat Alistair²^ORCID,Culpepper J. Shane¹^ORCID

Affiliation:

1. RMIT University, Melbourne, Australia

2. The University of Melbourne, Melbourne, Australia

Abstract

Rank fusion is a powerful technique that allows multiple sources of information to be combined into a single result set. Query variations covering the same information need represent one way in which different sources of information might arise. However, when implemented in the obvious manner, fusion over query variations is not cost-effective, at odds with the usual web-search requirement for strict per-query efficiency guarantees. In this work, we propose a novel solution to query fusion by splitting the computation into two parts: one phase that is carried out offline, to generate pre-computed centroid answers for queries addressing broadly similar information needs, and then a second online phase that uses the corresponding topic centroid to compute a result page for each query. To achieve this, we make use of score-based fusion algorithms whose costs can be amortized via the pre-processing step and that can then be efficiently combined during subsequent per-query re-ranking operations. Experimental results using the ClueWeb12B collection and the UQV100 query variations demonstrate that centroid-based approaches allow improved retrieval effectiveness at little or no loss in query throughput or latency and within reasonable pre-processing requirements. We additionally show that queries that do not match any of the pre-computed clusters can be accurately identified and efficiently processed in our proposed ranking pipeline.

Funder

Australian Research Training Program Scholarship

Australian Research Council’s Discovery Projects Scheme

RMIT Vice Chancellors PhD Scholarship

Amazon Research Award

Google Faculty Research Award

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Science Applications,General Business, Management and Accounting,Information Systems

Link

https://dl.acm.org/doi/pdf/10.1145/3345001

Reference92 articles.

1. Probabilistic models of information retrieval based on measuring the divergence from randomness

2. Design trade-offs for search engine caching

Cited by 22 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Is your search query well-formed? A natural query understanding for patent prior art search;World Patent Information;2024-03

2. Performance prediction of multivariable linear regression based on the optimal influencing factors for ranking aggregation in crowdsourcing task;Data Technologies and Applications;2023-07-04

3. Searching Parameterized Retrieval & Verification Loss for Re-Identification;IEEE Journal of Selected Topics in Signal Processing;2023-05

4. Improving Content Retrievability in Search with Controllable Query Generation;Proceedings of the ACM Web Conference 2023;2023-04-30

5. How do Human and Contextual Factors Affect the Way People Formulate Queries?;Proceedings of the 2023 Conference on Human Information Interaction and Retrieval;2023-03-19