Correcting model misspecification in relationship estimates

Author:

Jewett Ethan M.ORCID,

Abstract

1.AbstractThe datasets of large genotyping biobanks and direct-to-consumer genetic testing companies contain many related individuals. Until now, it has been widely accepted that the most distant relationships that can be detected are around fifteen degrees (approximately 8thcousins) and that practical relationship estimates have a ceiling around ten degrees (approximately 5thcousins). However, we show that these assumptions are incorrect and that they are due to a misapplication of relationship estimators. In particular, relationship estimators are applied almost exclusively to putative relatives who have been identified because they share detectable tracts of DNA identically by descent (IBD). However, no existing relationship estimator conditions on the event that two individuals share at least one detectable segment of IBD anywhere in the genome. As a result, the relationship estimates obtained using existing estimators are dramatically biased for distant relationships, inferring all sufficiently distant relationships to be around ten degrees regardless of the depth of the true relationship. Existing relationship estimators are derived under a model that assumes that each pair of related individuals shares a single common ancestor (or mating pair of ancestors). This model breaks down for relationships beyond 10 generations in the past because individuals share many thousands of cryptic common ancestors due to pedigree collapse. We first derive a corrected likelihood that conditions on the event that at least one segment is observed between a pair of putative relatives and we demonstrate that the corrected likelihood largely eliminates the bias in estimates of pairwise relationships and provides a more accurate characterization of the uncertainty in these estimates. We then reformulate the relationship inference problem to account for the fact that individuals share many common ancestors, not just one. We demonstrate that the most distant relationship that can be inferred using IBD may be 200 degrees or more, rather than ten, extending the time-to-common ancestor from approximately 300 years in the past to approximately 3,000 years in the past or more. This dramatic increase in the range of relationship estimators makes it possible to infer relationships whose common ancestors lived before historical events such as European settlement of the Americas, the Transatlantic Slave Trade, and the rise and fall of the Roman Empire.

Publisher

Cold Spring Harbor Laboratory

Reference22 articles.

1. C.A. Ball , M.J. Barber , J. Byrnes , P. Carbonetto , K.G. Chahine , R.E. Curtis , J.M. Granka , E. Han , E.L. Hong , A.R. Kermany , N.M. Myres , K. Noto , J. Qi , K. Rand , Y. Wang , and L. Willmore . Rapid forward-in-time simulation at the chromosome and genome level. https://www.ancestry.com/dna/resource/whitePaper/AncestryDNA-Matching-White-Paper.pdf, 2016.

2. Crossover interference and sex-specific genetic maps shape identical by descent sharing in close relatives

3. A novel single-gamma approximation to the sum of independent gamma variables, and a generalization to infinitely divisible distributions

4. Addressing the feasibility of people of african descent finding living african relatives using direct-to-consumer genetic testing;American Journal of Biological Anthropology,2023

5. Supporting the use of genetic genealogy in restoring family narratives following the transatlantic slave trade;Am Anthropol,2024

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3