A statistical boosting framework for polygenic risk scores based on large-scale genotype data-Reference-Cited by-同舟云学术

A statistical boosting framework for polygenic risk scores based on large-scale genotype data

Published:2023-01-10 Issue: Volume:13 Page:
ISSN:1664-8021
Container-title:Frontiers in Genetics
language:
Short-container-title:Front. Genet.

Author:

Klinkhammer Hannah,Staerk Christian,Maj Carlo,Krawitz Peter Michael,Mayr Andreas

Abstract

Polygenic risk scores (PRS) evaluate the individual genetic liability to a certain trait and are expected to play an increasingly important role in clinical risk stratification. Most often, PRS are estimated based on summary statistics of univariate effects derived from genome-wide association studies. To improve the predictive performance of PRS, it is desirable to fit multivariable models directly on the genetic data. Due to the large and high-dimensional data, a direct application of existing methods is often not feasible and new efficient algorithms are required to overcome the computational burden regarding efficiency and memory demands. We develop an adapted component-wise L2-boosting algorithm to fit genotype data from large cohort studies to continuous outcomes using linear base-learners for the genetic variants. Similar to the snpnet approach implementing lasso regression, the proposed snpboost approach iteratively works on smaller batches of variants. By restricting the set of possible base-learners in each boosting step to variants most correlated with the residuals from previous iterations, the computational efficiency can be substantially increased without losing prediction accuracy. Furthermore, for large-scale data based on various traits from the UK Biobank we show that our method yields competitive prediction accuracy and computational efficiency compared to the snpnet approach and further commonly used methods. Due to the modular structure of boosting, our framework can be further extended to construct PRS for different outcome data and effect types—we illustrate this for the prediction of binary traits.

Funder

Deutsche Forschungsgemeinschaft

Publisher

Frontiers Media SA

Subject

Genetics (clinical),Genetics,Molecular Medicine

Reference66 articles.

1. Blood pressure and human genetic variation in the general population;Arora;Curr. Opin. Cardiol.,2010

2. The emerging landscape of health research based on biobanks linked to electronic health records: Existing resources, statistical challenges, and potential opportunities;Beesley;Statistics Med.,2020

3. Boosting algorithms: Regularization, prediction and model fitting;Bühlmann;Stat. Sci.,2007

4. Boosting with the l2 loss;Bühlmann;J. Am. Stat. Assoc.,2003

5. Sparsity oracle inequalities for the Lasso;Bunea;Electron. J. Statistics,2007

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Boosting multivariate structured additive distributional regression models;Statistics in Medicine;2023-03-17

2. On the role of benchmarking data sets and simulations in method comparison studies;Biometrical Journal;2023-02-21