Genotype prediction of 336,463 samples from public expression data-Reference-Cited by-同舟云学术

Genotype prediction of 336,463 samples from public expression data

Published:2023-10-22 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Razi Afrooz^ORCID,Lo Christopher C.,Wang Siruo,Leek Jeffrey T.^ORCID,Hansen Kasper D.^ORCID

Abstract

AbstractTens of thousands of RNA-sequencing experiments comprising hundreds of thousands of individual samples have now been performed. These data represent a broad range of experimental conditions, sequencing technologies, and hypotheses under study. The Recount project has aggregated and uniformly processed hundreds of thousands of publicly available RNA-seq samples. Most of these samples only include RNA expression measurements; genotype data for these same samples would enable a wide range of analyses including variant prioritization, eQTL analysis, and studies of allele specific expression. Here, we developed a statistical model based on the existing reference and alternative read counts from the RNA-seq experiments available through Recount3 to predict genotypes at autosomal biallelic loci in coding regions. We demonstrate the accuracy of our model using large-scale studies that measured both gene expression and genotype genome-wide. We show that our predictive model is highly accurate with 99.5% overall accuracy, 99.6% major allele accuracy, and 90.4% minor allele accuracy. Our model is robust to tissue and study effects, provided the coverage is high enough. We applied this model to genotype all the samples in Recount3 and provide the largest ready-to-use expression repository containing genotype information. We illustrate that the predicted genotype from RNA-seq data is sufficient to unravel the underlying population structure of samples in Recount3 using Principal Component Analysis.

Publisher

Cold Spring Harbor Laboratory

Reference34 articles.

1. Variant analysis pipeline for accurate detection of genomic variants from transcriptome sequencing data

2. RNA-Seq based genetic variant discovery provides new insights into controlling fat deposition in the tail of sheep

3. Ancestry patterns inferred from massive RNA-seq data

4. GenBank

5. Quantifying uncertainty in genotype calls

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Large-scale genotype prediction from RNA sequence data necessitates a new ethical and policy framework;Nature Genetics;2024-07-22