Overcoming Long Inference Time of Nearest Neighbors Analysis in Regression and Uncertainty Prediction-Reference-Cited by-同舟云学术

Overcoming Long Inference Time of Nearest Neighbors Analysis in Regression and Uncertainty Prediction

Published:2024-04-24 Issue:5 Volume:5 Page:
ISSN:2661-8907
Container-title:SN Computer Science
language:en
Short-container-title:SN COMPUT. SCI.

Author:

Koutenský František,Šimánek Petr,Čepek Miroslav,Kovalenko Alexander^ORCID

Abstract

AbstractThe intuitive approach of comparing like with like, forms the basis of the so-called nearest neighbor analysis, which is central to many machine learning algorithms. Nearest neighbor analysis is easy to interpret, analyze, and reason about. It is widely used in advanced techniques such as uncertainty estimation in regression models, as well as the renowned k-nearest neighbor-based algorithms. Nevertheless, its high inference time complexity, which is dataset size dependent even in the case of its faster approximated version, restricts its applications and can considerably inflate the application cost. In this paper, we address the problem of high inference time complexity. By using gradient-boosted regression trees as a predictor of the labels obtained from nearest neighbor analysis, we demonstrate a significant increase in inference speed, improving by several orders of magnitude. We validate the effectiveness of our approach on a real-world European Car Pricing Dataset with approximately

$$4.2 \times 10^6$$

4.2 × 10 6 rows for both residual cost and price uncertainty prediction. Moreover, we assess our method’s performance on the most commonly used tabular benchmark datasets to demonstrate its scalability. The link is to github repository where the code is available: https://github.com/koutefra/uncertainty_experiments.

Funder

Technologická Agentura České Republiky

Czech Technical University in Prague

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1007/s42979-024-02670-2.pdf

Reference34 articles.

1. Bhatia N, et al. Survey of nearest neighbor techniques. arXiv preprint arXiv:1007.0085. 2010

2. Zafar MR, Khan N. Deterministic local interpretable model-agnostic explanations for stable explainability. Mach Learn Knowl Extr. 2021;3(3):525–41.

3. Winter B, Matlock T. Making judgments based on similarity and proximity. Metaphor Symb. 2013;28(4):219–32.

4. Fix E, Hodges J. An important contribution to nonparametric discriminant analysis and density estimation. Int Stat Rev. 1951;3(57):233–8.

5. Mukhlishin MF, Saputra R, Wibowo A, Predicting house sale price using fuzzy logic, artificial neural network and k-nearest neighbor. In: 2017 1st international conference on informatics and computational sciences (ICICoS). IEEE, 2017; p. 171–6.