Kennard-Stone method outperforms the Random Sampling in the selection of calibration samples in SNPs and NIR data-Reference-Cited by-同舟云学术

Kennard-Stone method outperforms the Random Sampling in the selection of calibration samples in SNPs and NIR data

Published:2022 Issue:5 Volume:52 Page:
ISSN:1678-4596
Container-title:Ciência Rural
language:
Short-container-title:Cienc. Rural

Author:

Ferreira Roberta de Amorim¹^ORCID,Teixeira Gabriely²^ORCID,Peternelli Luiz Alexandre²^ORCID

Affiliation:

1. Universidade Federal de Viçosa (UFV),, Brazil; Instituto Federal de Minas Gerais (IFMG), Brazil

2. Universidade Federal de Viçosa (UFV),, Brazil

Abstract

ABSTRACT: Splitting the whole dataset into training and testing subsets is a crucial part of optimizing models. This study evaluated the influence of the choice of the training subset in the construction of predictive models, as well as on their validation. For this purpose we assessed the Kennard-Stone (KS) and the Random Sampling (RS) methods in near-infrared spectroscopy data (NIR) and marker data SNPs (Single Nucleotide Polymorphisms). It is worth noting that in SNPs data, there is no knowledge of reports in the literature regarding the use of the KS method. For the construction and validation of the models, the partial least squares (PLS) estimation method and the Bayesian Lasso (BLASSO) proved to be more efficient for NIR data and for marker data SNPs, respectively. The evaluation of the predictive capacity of the models obtained after the data partition occurred through the correlation between the predicted and the observed values, and the corresponding square root of the mean squared error of prediction. For both datasets, results indicated that the results from KS and RS methods differ statistically from each other by the F test (P-value < 0.01). The KS method showed to be more efficient than RS in practically all repetitions. Also, KS method has the advantage of being easy and fast to be applied and also to select the same samples, which provides excellent benefits in the following analyses.

Publisher

FapUNIFESP (SciELO)

Subject

General Veterinary,Agronomy and Crop Science,Animal Science and Zoology

Reference46 articles.

1. Optimization of genomic selection training populations with a genetic algorithm.;AKDEMIR D.;Genetics Selection Evolution,2015

2. Prediction of lignin content in Different Parts of Sugarcane Using Near-Infrared Spectroscopy (NIR), Ordered Predictors Selection (OPS), and Partial Least Squares (PLS).;ASSIS C.;Applied Spectroscopy,2017

3. Independent component regression applied to genomic selection for carcass traits in pigs;AZEVEDO C.;Pesquisa Agropecuaria Brasileira,2013

4. Elementos de Amostragem;BOLFARINE H.,2005

5. Chemical Systems Under Indirect Observation: Latent Properties and Chemometrics.;BROWN S.;Applied Spectroscopy,1995

Cited by 14 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Comparing the potential of benchtop and handheld mid-infrared spectrometers for predicting soil phosphorus (P) sorption capacity and evaluating the influence of sample preparation;Spectrochimica Acta Part A: Molecular and Biomolecular Spectroscopy;2024-12

2. Advanced chemometrics toward robust spectral analysis for fruit quality evaluation;Trends in Food Science & Technology;2024-08

3. Improve the accuracy of FT-NIR for determination of zearalenone content in wheat by using the characteristic wavelength optimization algorithm;Spectrochimica Acta Part A: Molecular and Biomolecular Spectroscopy;2024-05

4. SVR Chemometrics to Quantify β-Lactoglobulin and α-Lactalbumin in Milk Using MIR;Foods;2024-01-03

5. A deep spectral prediction network to quantitatively determine heavy metal elements in soil by X-ray fluorescence;Journal of Analytical Atomic Spectrometry;2024