Cache-Efficient Top-k Aggregation over High Cardinality Large Datasets-Reference-Cited by-同舟云学术

Cache-Efficient Top-k Aggregation over High Cardinality Large Datasets

Published:2023-12 Issue:4 Volume:17 Page:644-656
ISSN:2150-8097
Container-title:Proceedings of the VLDB Endowment
language:en
Short-container-title:Proc. VLDB Endow.

Author:

Siddiqui Tarique¹,Narasayya Vivek¹,Dumitru Marius²,Chaudhuri Surajit¹

Affiliation:

1. Microsoft Research, Redmond, Washington, USA

2. Microsoft, Redmond, Washington, USA

Abstract

Top-k aggregation queries are widely used in data analytics for summarizing and identifying important groups from large amounts of data. These queries are usually processed by first computing exact aggregates for all groups and then selecting the groups with the top-k aggregate values. However, such an approach can be inefficient for high-cardinality large datasets where intermediate results may not fit within the local cache of multi-core processors leading to excessive data movement. To address this problem, we have developed Zippy, a new cache-conscious aggregation framework that leverages the skew in the data distribution to minimize data movements. This is achieved by designing cache-resident data structures and an adaptive multi-pass algorithm that quickly identifies candidate groups during processing, and performs exact aggregations for these groups. The non-candidate groups are pruned cheaply using efficient hashing and partitioning techniques without performing exact aggregations. We develop techniques to improve robustness over adversarial data distributions and have optimized the framework to reuse computations incrementally for rolling (or paginated) top-k aggregate queries. Our extensive evaluation using both real-world and synthetic datasets demonstrate that Zippy can achieve a median speed-up of more than 3× for monotonic aggregation functions across typical ranges of k values (e.g., 1 to 100) and 1.4× for non-monotonic functions when compared with state-of-the-art cache-conscious aggregation techniques.

Publisher

Association for Computing Machinery (ACM)

Link

https://dl.acm.org/doi/pdf/10.14778/3636218.3636222

Reference35 articles.

1. 2023. Apache Datafusion). https://godatadriven.com/blog/optimizing-topk-queries-in-datafusion/ [Online; accessed 3-May-2023].

2. 2023. PowerBI (https://powerbi.microsoft.com/en-us/). https://powerbi.microsoft.com/en-us/ [Online; accessed 3-May-2023].

3. 2023. Tableau Public (www.tableaupublic.com/). www.tableaupublic.com/ [Online; accessed 3-May-2023].

4. Martina-Cezara Albutiu, Alfons Kemper, and Thomas Neumann. 2012. Massively parallel sort-merge joins in main memory multi-core database systems. arXiv preprint arXiv:1207.0145 (2012).

5. Multi-core, main-memory joins