An Approximate Algorithm for Maximum Inner Product Search over Streaming Sparse Vectors-Reference-Cited by-同舟云学术

An Approximate Algorithm for Maximum Inner Product Search over Streaming Sparse Vectors

Published:2023-11-08 Issue:2 Volume:42 Page:1-43
ISSN:1046-8188
Container-title:ACM Transactions on Information Systems
language:en
Short-container-title:ACM Trans. Inf. Syst.

Author:

Bruch Sebastian¹^ORCID,Nardini Franco Maria²^ORCID,Ingber Amir³^ORCID,Liberty Edo¹^ORCID

Affiliation:

1. Pinecone, USA

2. ISTI-CNR, Italy

3. Pinecone, Israel

Abstract

Maximum Inner Product Search or top- k retrieval on sparse vectors is well understood in information retrieval, with a number of mature algorithms that solve it exactly. However, all existing algorithms are tailored to text and frequency-based similarity measures. To achieve optimal memory footprint and query latency, they rely on the near stationarity of documents and on laws governing natural languages. We consider, instead, a setup in which collections are streaming—necessitating dynamic indexing—and where indexing and retrieval must work with arbitrarily distributed real-valued vectors. As we show, existing algorithms are no longer competitive in this setup, even against naïve solutions. We investigate this gap and present a novel approximate solution, called Sinnamon , that can efficiently retrieve the top- k results for sparse real valued vectors drawn from arbitrary distributions. Notably, Sinnamon offers levers to trade off memory consumption, latency, and accuracy, making the algorithm suitable for constrained applications and systems. We give theoretical results on the error introduced by the approximate nature of the algorithm and present an empirical evaluation of its performance on two hardware platforms and synthetic and real-valued datasets. We conclude by laying out concrete directions for future research on this general top- k retrieval problem over sparse vectors.

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Science Applications,General Business, Management and Accounting,Information Systems

Link

https://dl.acm.org/doi/pdf/10.1145/3609797

Reference67 articles.

1. The Fast Johnson–Lindenstrauss Transform and Approximate Nearest Neighbors

2. An almost optimal unrestricted fast Johnson-Lindenstrauss transform;Ailon Nir;ACM Trans. Algor.,2013

3. Nima Asadi. 2013. Multi-stage Search Architectures for Streaming Documents. University of Maryland.

4. Fast candidate generation for two-phase document ranking

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Efficient Approximate Maximum Inner Product Search Over Sparse Vectors;2024 IEEE 40th International Conference on Data Engineering (ICDE);2024-05-13

2. Two-Step SPLADE: Simple, Efficient and Effective Approximation of SPLADE;Lecture Notes in Computer Science;2024