Algorithmic Stability and Sanity-Check Bounds for Leave-One-Out Cross-Validation-Reference-Cited by-同舟云学术

Algorithmic Stability and Sanity-Check Bounds for Leave-One-Out Cross-Validation

Published:1999-08-01 Issue:6 Volume:11 Page:1427-1453
ISSN:0899-7667
Container-title:Neural Computation
language:en
Short-container-title:Neural Computation

Author:

Kearns Michael¹,Ron Dana²

Affiliation:

1. AT&T Labs Research, Florham Park, NJ 07932, U.S.A.

2. Department of EE—Systems, Tel Aviv University, 69978 Ramat Aviv, Israel

Abstract

In this article we prove sanity-check bounds for the error of the leave-oneout cross-validation estimate of the generalization error: that is, bounds showing that the worst-case error of this estimate is not much worse than that of the training error estimate. The name sanity check refers to the fact that although we often expect the leave-one-out estimate to perform considerably better than the training error estimate, we are here only seeking assurance that its performance will not be considerably worse. Perhaps surprisingly, such assurance has been given only for limited cases in the prior literature on cross-validation. Any nontrivial bound on the error of leave-one-out must rely on some notion of algorithmic stability. Previous bounds relied on the rather strong notion of hypothesis stability, whose application was primarily limited to nearest-neighbor and other local algorithms. Here we introduce the new and weaker notion of error stability and apply it to obtain sanity-check bounds for leave-one-out for other classes of learning algorithms, including training error minimization procedures and Bayesian algorithms. We also provide lower bounds demonstrating the necessity of some form of error stability for proving bounds on the error of the leave-one-out estimate, and the fact that for training error minimization algorithms, in the worst case such bounds must still depend on the Vapnik-Chervonenkis dimension of the hypothesis class.

Publisher

MIT Press - Journals

Subject

Cognitive Neuroscience,Arts and Humanities (miscellaneous)

Link

https://www.mitpressjournals.org/doi/pdf/10.1162/089976699300016304

Reference8 articles.

1. Distribution-free inequalities for the deleted and holdout error estimates

2. Distribution-free performance bounds for potential function rules

3. Stochastic Relaxation, Gibbs Distributions, and the Bayesian Restoration of Images

4. Decision theoretic generalizations of the PAC model for neural net and other learning applications

Cited by 290 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Cross-validation on extreme regions;Extremes;2024-09-03

2. Enabling Reliable Visual Detection of Chronic Myocardial Infarction with Native T1 Cardiac MRI Using Data-Driven Native Contrast Mapping;Radiology: Cardiothoracic Imaging;2024-08-01

3. Fast, Robust and Interpretable Participant Contribution Estimation for Federated Learning;2024 IEEE 40th International Conference on Data Engineering (ICDE);2024-05-13

4. Generalization bounds for learning under graph-dependence: a survey;Machine Learning;2024-04-03

5. CasMDN: A deep learning-based multivariate distribution modelling approach and its application in geotechnical engineering;Computers and Geotechnics;2024-04