New wine in an old bottle-Reference-Cited by-同舟云学术

New wine in an old bottle

Published:2022-05 Issue:9 Volume:15 Page:1924-1936
ISSN:2150-8097
Container-title:Proceedings of the VLDB Endowment
language:en
Short-container-title:Proc. VLDB Endow.

Author:

Bhattacharya Arindam¹,Gudesa Chathur²,Bagchi Amitabha¹,Bedathur Srikanta¹

Affiliation:

1. CSE, IIT Delhi, India

2. EE, IIT Delhi, India

Abstract

In many applications of Bloom filters, it is possible to exploit the patterns present in the inserted and non-inserted keys to achieve more compression than the standard Bloom filter. A new class of Bloom filters called Learned Bloom filters use machine learning models to exploit these patterns in the data. In practice, these methods and their variants raise many questions: the choice of machine learning models, the training paradigm to achieve the desired results, the choice of thresholds, the number of partitions in case multiple partitions are used, and other such design decisions. In this paper, we present a simple partitioned Bloom filter that works as follows: we partition the Bloom filter into segments, each of which uses a simple projection-based hash function computed using the data. We also provide a theoretical analysis that provides a principled way to select the design parameters of our method: number of hash functions and number of bits per partition. We perform empirical evaluations of our methods on various real-world datasets spanning several applications. We show that it can achieve an improvement in false positive rates of up to two orders of magnitude over standard Bloom filters for the same memory usage, and upto 50% better compression (bytes used per key) for same FPR, and, consistently beats the existing variants of learned Bloom filters.

Publisher

Association for Computing Machinery (ACM)

Subject

General Earth and Planetary Sciences,Water Science and Technology,Geography, Planning and Development

Link

https://dl.acm.org/doi/pdf/10.14778/3538598.3538613

Reference29 articles.

1. Nir Ailon and Bernard Chazelle. 2006. Approximate nearest neighbors and the fast Johnson-Lindenstrauss transform. In SoTC. Nir Ailon and Bernard Chazelle. 2006. Approximate nearest neighbors and the fast Johnson-Lindenstrauss transform. In SoTC.

2. Hyrum S Anderson and Phil Roth . 2018. Ember: an open dataset for training static PE malware machine learning models. arXiv preprint arXiv:1804.04637 ( 2018 ). Hyrum S Anderson and Phil Roth. 2018. Ember: an open dataset for training static PE malware machine learning models. arXiv preprint arXiv:1804.04637 (2018).

3. Searching for exotic particles in high-energy physics with deep learning

4. Arindam Bhattacharya Srikanta Bedathur and Amitabha Bagchi. 2020. Adaptive Learned Bloom Filters under Incremental Workloads. In CoDS-COMAD. Arindam Bhattacharya Srikanta Bedathur and Amitabha Bagchi. 2020. Adaptive Learned Bloom Filters under Incremental Workloads. In CoDS-COMAD.

5. Arindam Bhattacharya Sumanth Varambally Amitabha Bagchi and Srikanta Bedathur. 2021. Fast One-class Classification using Class Boundary-preserving Random Projections. In KDD. Arindam Bhattacharya Sumanth Varambally Amitabha Bagchi and Srikanta Bedathur. 2021. Fast One-class Classification using Class Boundary-preserving Random Projections. In KDD.

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. The Reinforcement Cuckoo Filter;IEEE INFOCOM 2024 - IEEE Conference on Computer Communications;2024-05-20

2. A novel revocation management for distributed environment: a detailed study;Cluster Computing;2023-08-26

3. PA-LBF: Prefix-Based and Adaptive Learned Bloom Filter for Spatial Data;International Journal of Intelligent Systems;2023-03-29