Abstract
AbstractInferring the frequency and mode of hybridization among closely related organisms is an important step for understanding the process of speciation and can help to uncover reticulated patterns of phylogeny more generally. Phylogenomic methods to test for the presence of hybridization come in many varieties and typically operate by leveraging expected patterns of genealogical discordance in the absence of hybridization. An important assumption made by these tests is that the data (genes or SNPs) are independent given the species tree. However, when the data are closely linked, it is especially important to consider their non-independence. Recently, deep learning techniques such as convolutional neural networks (CNNs) have been used to perform population genetic inferences with linked SNPs coded as binary images. Here we use CNNs for selecting among candidate hybridization scenarios using the tree topology (((P1,P2),P3),Out) and a matrix of pairwise nucleotide divergence (dXY) calculated in windows across the genome. Using coalescent simulations to train and independently test a neural network showed that our method, HyDe-CNN, was able to accurately perform model selection for hybridization scenarios across a wide-breath of parameter space. We then used HyDe-CNN to test models of admixture inHeliconiusbutterflies, as well as comparing it to a random forest classifier trained on introgression-based statistics. Given the flexibility of our approach, the dropping cost of long-read sequencing, and the continued improvement of CNN architectures, we anticipate that inferences of hybridization using deep learning methods like ours will help researchers to better understand patterns of admixture in their study organisms.
Publisher
Cold Spring Harbor Laboratory
Reference75 articles.
1. Abadi M , Agarwal A , Barham P , et al. (2016) TensorFlow: Large-scale machine learning on heterogeneous distributed systems. CoRR, abs/1603.04467.
2. Predicting the landscape of recombination using deep learning;Molecular Biology and Evolution,2020
3. Agarap AF (2018) Deep learning using rectified linear units (ReLU). CoRR, abs/1803.08375.
4. Opportunities and challenges in long-read sequencing data analysis;Genome Biology,2020
5. Anderson E (1949) Introgressive hybridization. John Wiley, New York, NY, USA.
Cited by
5 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献