Random sampling for histogram construction-Reference-Cited by-同舟云学术

Random sampling for histogram construction

Published:1998-06 Issue:2 Volume:27 Page:436-447
ISSN:0163-5808
Container-title:ACM SIGMOD Record
language:en
Short-container-title:SIGMOD Rec.

Author:

Chaudhuri Surajit¹,Motwani Rajeev²,Narasayya Vivek¹

Affiliation:

1. Microsoft Research

2. Stanford University

Abstract

Random sampling is a standard technique for constructing (approximate) histograms for query optimization. However, any real implementation in commercial products requires solving the hard problem of determining “How much sampling is enough?” We address this critical question in the context of equi-height histograms used in many commercial products, including Microsoft SQL Server. We introduce a conservative error metric capturing the intuition that for an approximate histogram to have low error, the error must be small in all regions of the histogram. We then present a result establishing an optimal bound on the amount of sampling required for pre-specified error bounds. We also describe an adaptive page sampling algorithm which achieves greater efficiency by using all values in a sampled page but adjusts the amount of sampling depending on clustering of values in pages. Next, we establish that the problem of estimating the number of distinct values is provably difficult , but propose a new error metric which has a reliable estimator and can still be exploited by query optimizers to influence the choice of execution plans. The algorithm for histogram construction was prototyped on Microsoft SQL Server 7.0 and we present experimental results showing that the adaptive algorithm accurately approximates the true histogram over different data distributions.

Publisher

Association for Computing Machinery (ACM)

Subject

Information Systems,Software

Link

https://dl.acm.org/doi/pdf/10.1145/276305.276343

Reference29 articles.

1. Estimating the Number of Species: A Review

2. Estimation of the size of a closed population when capture probabilities vary among animals

3. Robust Estimation of Population Size When Capture Probabilities Vary Among Animals

4. A. Chao. Nonparametric estimation of the number of classes in a population. Scandinavian Journal o/Statistical Theory and Applications 11(1984): 265-270. A. Chao. Nonparametric estimation of the number of classes in a population. Scandinavian Journal o/Statistical Theory and Applications 11(1984): 265-270.

5. S. Chaudhuri R. Motwani and V. Narasayya. Using Random Sampling for Histogram Construction. Microsoft Research Report In preparation 1997. S. Chaudhuri R. Motwani and V. Narasayya. Using Random Sampling for Histogram Construction. Microsoft Research Report In preparation 1997.

Cited by 66 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. CDFRS: A scalable sampling approach for efficient big data analysis;Information Processing & Management;2024-07

2. Testing Closeness of Multivariate Distributions via Ramsey Theory;Proceedings of the 56th Annual ACM Symposium on Theory of Computing;2024-06-10

3. ROME: Robust Query Optimization via Parallel Multi-Plan Execution;Proceedings of the ACM on Management of Data;2024-05-29

4. Efficient Random Sampling from Very Large Databases;Lecture Notes in Computer Science;2024

5. On distributed data aggregation and the precision of approximate histograms;Journal of Parallel and Distributed Computing;2023-10