Sizing sketches-Reference-Cited by-同舟云学术

Sizing sketches

Published:2007-06-12 Issue:1 Volume:35 Page:157-168
ISSN:0163-5999
Container-title:ACM SIGMETRICS Performance Evaluation Review
language:en
Short-container-title:SIGMETRICS Perform. Eval. Rev.

Author:

Wang Zhe¹,Dong Wei¹,Josephson William¹,Lv Qin¹,Charikar Moses¹,Li Kai¹

Affiliation:

1. Princeton University, Princeton, NJ

Abstract

Sketches are compact data structures that can be used to estimate properties of the original data in building large-scale search engines and data analysis systems. Recent theoretical and experimental studies have shown that sketches constructed from feature vectors using randomized projections can effectively approximate L1 distance on the feature vectors with the Hamming distance on their sketches. Furthermore, such sketches can achieve good filtering accuracy while reducing the metadata space requirement and speeding up similarity searches by an order of magnitude. However, it is not clear how to choose the size of the sketches since it depends ondata type, dataset size, and desired filtering quality. In real systems designs, it is necessary to understand how to choose sketch size without the dataset, or at least without the whole datase. This paper presents an analytical model and experimental results to help system designers make such design decisions. We present arank-based filtering model that describes the relationship between sketch size and data set size based on the dataset distance distribution. Our experimental results with several datasets including images, audio, and 3D shapes show that the model yields good, conservative predictions. We show that the parameters of the model can be set with a small sample data set and the resulting model can make good predictions for a large dataset. We illustrate how to apply the approach with a concrete example.

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Networks and Communications,Hardware and Architecture,Software

Link

https://dl.acm.org/doi/pdf/10.1145/1269899.1254900

Reference24 articles.

1. S. Balko I. Schmitt and G. Saake. The active vertice method: A performance filtering approach to high-dimensional indexing. Elsevier Data and Knowledge Engineering (DKE) 51(3):369--397 2004. 10.1016/j.datak.2004.06.002 S. Balko I. Schmitt and G. Saake. The active vertice method: A performance filtering approach to high-dimensional indexing. Elsevier Data and Knowledge Engineering (DKE) 51(3):369--397 2004. 10.1016/j.datak.2004.06.002

2. K-d trees for semidynamic point sets

3. Cover trees for nearest neighbor

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Linear time identification of local and global outliers;Neurocomputing;2021-03

2. Ensemble dimensionality reduction and feature gene extraction for single-cell RNA-seq data;Nature Communications;2020-11-17

3. Pivot Selection for Narrow Sketches by Optimization Algorithms;Similarity Search and Applications;2020

4. Fast Filtering for Nearest Neighbor Search by Sketch Enumeration Without Using Matching;AI 2019: Advances in Artificial Intelligence;2019

5. Discovery of rare cells from voluminous single cell expression data;Nature Communications;2018-11-09