Identifying the best approximating model in Bayesian phylogenetics: Bayes factors, cross-validation or wAIC?-Reference-Cited by-同舟云学术

Identifying the best approximating model in Bayesian phylogenetics: Bayes factors, cross-validation or wAIC?

Published:2022-04-22 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Lartillot Nicolas^ORCID

Abstract

AbstractThere is still no consensus as to how to select models in Bayesian phylogenetics, and more generally in applied Bayesian statistics. Bayes factors are often presented as the method of choice, yet other approaches have been proposed, such as cross-validation or information criteria. Each of these paradigms raises specific computational challenges, but they also differ in their statistical meaning, being motivated by different objectives: either testing hypotheses or finding the best-approximating model. These alternative goals entail different compromises, and as a result, Bayes factors, cross-validation and information criteria may be valid for addressing different questions. Here, the question of Bayesian model selection is revisited, with a focus on the problem of finding the best-approximating model. Several model selection approaches were re-implemented, numerically assessed and compared: Bayes factors, cross-validation (CV), in its different forms (k-fold or leave-one-out), and the widely applicable information criterion (wAIC), which is asymptotically equivalent to leave-one-out cross validation (LOO-CV). Using a combination of analytical results and empirical and simulation analyses, it is shown that Bayes factors are unduly conservative. In contrast, cross-validation represents a more adequate formalism for selecting the model returning the best approximation of the data-generating process and the most accurate estimates of the parameters of interest. Among alternative CV schemes, LOO-CV and its asymptotic equivalent represented by the wAIC, stand out as the best choices, conceptually and computationally, given that both can be simultaneously computed based on standard MCMC runs under the posterior distribution.

Publisher

Cold Spring Harbor Laboratory

Reference88 articles.

1. Model selection for ecologists: the worldviews of AIC and BIC

2. A new look at the statistical model identification;IEEE Trans. Automat. Contr.,1974

3. Bayesian evolutionary model testing in the phylogenomics era: matching model complexity with computational efficiency

4. Improving the Accuracy of Demographic and Molecular Clock Model Comparison While Accommodating Phylogenetic Uncertainty

5. Make the most of your samples: Bayes factor estimators for high-dimensional models of sequence evolution

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Detecting episodic evolution through Bayesian inference of molecular clock models;2023-06-19

2. Improved modelling of compositional heterogeneity reconciles phylogenomic conflicts among lacewings;Palaeoentomology;2023-02-28