Use of the linear regression method to evaluate population accuracy of predictions from non-linear models-Reference-Cited by-同舟云学术

Use of the linear regression method to evaluate population accuracy of predictions from non-linear models

Published:2024-05-31 Issue: Volume:15 Page:
ISSN:1664-8021
Container-title:Frontiers in Genetics
language:
Short-container-title:Front. Genet.

Author:

Yu Haipeng,Fernando Rohan L.,Dekkers Jack C. M.

Abstract

BackgroundTo address the limitations of commonly used cross-validation methods, the linear regression method (LR) was proposed to estimate population accuracy of predictions based on the implicit assumption that the fitted model is correct. This method also provides two statistics to determine the adequacy of the fitted model. The validity and behavior of the LR method have been provided and studied for linear predictions but not for nonlinear predictions. The objectives of this study were to 1) provide a mathematical proof for the validity of the LR method when predictions are based on conditional means, regardless of whether the predictions are linear or non-linear 2) investigate the ability of the LR method to detect whether the fitted model is adequate or inadequate, and 3) provide guidelines on how to appropriately partition the data into training and validation such that the LR method can identify an inadequate model.ResultsWe present a mathematical proof for the validity of the LR method to estimate population accuracy and to determine whether the fitted model is adequate or inadequate when the predictor is the conditional mean, which may be a non-linear function of the phenotype. Using three partitioning scenarios of simulated data, we show that the one of the LR statistics can detect an inadequate model only when the data are partitioned such that the values of relevant predictor variables differ between the training and validation sets. In contrast, we observed that the other LR statistic was able to detect an inadequate model for all three scenarios.ConclusionThe LR method has been proposed to address some limitations of the traditional approach of cross-validation in genetic evaluation. In this paper, we showed that the LR method is valid when the model is adequate and the conditional mean is the predictor, even when it is a non-linear function of the phenotype. We found one of the two LR statistics is superior because it was able to detect an inadequate model for all three partitioning scenarios (i.e., between animals, by age within animals, and between animals and by age) that were studied.

Publisher

Frontiers Media SA

Reference26 articles.

1. Correcting for base-population differences and unknown parent groups in single-step genomic predictions of Norwegian red cattle;Belay;J. Anim. Sci.,2022

2. Validation of single-step gblup genomic predictions from threshold models using the linear regression method: an application in chicken mortality;Bermann;J. Anim. Breed. Genet.,2021

3. Prospects for genomewide selection for quantitative traits in maize;Bernardo;Crop Sci.,2007

4. Julia: a fresh approach to numerical computing;Bezanson;SIAM Rev. Soc. Ind. Appl. Math.,2017

5. Modelling the variation in performance of a population of growing pig as affected by lysine supply and feeding strategy;Brossard;animal,2009