Optimisation of the core subset for the APY approximation of genomic relationships

Author:

Pocrnic IvanORCID,Lindgren FinnORCID,Tolhurst DanielORCID,Herring William O.,Gorjanc GregorORCID

Abstract

AbstractBackgroundBy entering the era of mega-scale genomics, we are facing many computational issues with standard genomic evaluation models due to their dense data structure and cubic computational complexity. Several scalable approaches have have been proposed to address this challenge, like the Algorithm for Proven and Young (APY). In APY, genotyped animals are partitioned into core and non-core subsets, which induces a sparser inverse of genomic relationship matrix. The partitioning into subsets is often done at random. While APY is a good approximation of the full model, the random partitioning can make results unstable, possibly affecting accuracy or even reranking animals. Here we present a stable optimisation of the core subset by choosing animals with the most informative genotype data.MethodsWe derived a novel algorithm for optimising the core subset based on the conditional genomic relationship matrix or the conditional SNP genotype matrix. We compared accuracy of genomic predictions with different core subsets on simulated and real pig data. The core subsets were constructed (1) at random, (2) based on the diagonal of genomic relationship matrix, (3) at random with weights from (2), or (4) based on the novel conditional algorithm. To understand the different core subset constructions, we have visualised population structure of genotyped animals with the linear Principal Component Analysis and the non-linear Uniform Manifold Approximation and Projection.ResultsAll core subset constructions performed equally well when the number of core animals captured most of variation in genomic relationships, both in simulated and real data. When the number of core animals was not optimal, there was substantial variability in results with the random construction and no variability with the conditional construction. Visualisation of population structure and chosen core animals showed that the conditional construction spreads core animals across the whole domain of genotyped animals in a repeatable manner.ConclusionsOur results confirm that the size of the core subset in APY is critical. The results further show that the core subset can be optimised with the conditional algorithm that achieves a good and repeatable spread of core animals across the domain of genotyped animals.

Publisher

Cold Spring Harbor Laboratory

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3