Using all gene families vastly expands data available for phylogenomic inference in primates-Reference-Cited by-同舟云学术

Using all gene families vastly expands data available for phylogenomic inference in primates

Published:2021-09-22 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Smith Megan L.^ORCID,Vanderpool Dan^ORCID,Hahn Matthew W.^ORCID

Abstract

AbstractTraditionally, single-copy orthologs have been the gold standard in phylogenomics. Most phylogenomic studies identify putative single-copy orthologs by using clustering approaches and retaining families with a single sequence from each species. However, this approach can severely limit the amount of data available by excluding larger families. Recent methodological advances have suggested several ways to include data from larger families. For instance, tree-based decomposition methods facilitate the extraction of orthologs from large families. Additionally, several popular methods for species tree inference appear to be robust to the inclusion of paralogs, and hence could use all of the data from larger families. Here, we explore the effects of using all families for phylogenetic inference using genomes from 26 primate species. We compare single-copy families, orthologs extracted using tree-based decomposition approaches, and all families with all data (i.e., including orthologs and paralogs). We explore several species tree inference methods, finding that across all nodes of the tree except one, identical trees are returned across nearly all datasets and methods. As in previous studies, the relationships among Platyrrhini remain contentious; however, the tree inference methods matter more than the dataset used. We also assess the effects of each dataset on branch length estimates, measures of phylogenetic uncertainty and concordance, and in detecting introgression. Our results demonstrate that using data from larger gene families drastically increases the number of genes available for phylogenetic inference and leads to consistent estimates of branch lengths, nodal certainty and concordance, and inferences of introgression.

Publisher

Cold Spring Harbor Laboratory

Reference60 articles.

1. Altenhoff AM , Glover NM , Dessimoz C. 2019. Inferring orthology and paralogy. In: Anisimova M , editor. Evolutionary genomics: Statistical and computational methods. New York, NY: Springer. p. 149–175. Available from: https://doi.org/10.1007/978-1-4939-9074-0_5

2. BLAST+: architecture and applications

3. trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses

4. Terrace Aware Data Structure for Phylogenomic Inference from Supermatrices

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. OrthoSNAP: A tree splitting and pruning algorithm for retrieving single-copy orthologs from gene family trees;PLOS Biology;2022-10-13

2. orthoSNAP: a tree splitting and pruning algorithm for retrieving single-copy orthologs from gene family trees;2021-11-02