A Fast Data-Driven Method for Genotype Imputation, Phasing, and Local Ancestry Inference: MendelImpute.jl-Reference-Cited by-同舟云学术

A Fast Data-Driven Method for Genotype Imputation, Phasing, and Local Ancestry Inference: MendelImpute.jl

Published:2020-10-25 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Chu Benjamin B.^ORCID,Sobel Eric M.,Wasiolek Rory,Sinsheimer Janet S.^ORCID,Zhou Hua^ORCID,Lange Kenneth

Abstract

1AbstractCurrent methods for genotype imputation and phasing exploit the sheer volume of data in haplotype reference panels and rely on hidden Markov models. Existing programs all have essentially the same imputation accuracy, are computationally intensive, and generally require pre-phasing the typed markers. We propose a novel data-mining method for genotype imputation and phasing that substitutes highly efficient linear algebra routines for hidden Markov model calculations. This strategy, embodied in our Julia program MendelImpute.jl, avoids explicit assumptions about recombination and population structure while delivering similar prediction accuracy, better memory usage, and an order of magnitude or better run-times compared to the fastest competing method. MendelImpute operates on both dosage data and unphased genotype data and simultaneously imputes missing genotypes and phase at both the typed and untyped SNPs. Finally, MendelImpute naturally extends to global and local ancestry estimation and lends itself to new strategies for data compression and hence faster data transport and sharing.

Publisher

Cold Spring Harbor Laboratory

Reference28 articles.

1. A global reference for human genetic variation

2. An integrated map of genetic variation from 1,092 human genomes

3. Fast model-based estimation of ancestry in unrelated individuals

4. Enhanced Methods for Local Ancestry Assignment in Sequenced Admixed Individuals