Knockoff boosted tree for model-free variable selection-Reference-Cited by-同舟云学术

Knockoff boosted tree for model-free variable selection

Published:2020-09-23 Issue:7 Volume:37 Page:976-983
ISSN:1367-4803
Container-title:Bioinformatics
language:en
Short-container-title:

Author:

Jiang Tao¹,Li Yuanyuan²,Motsinger-Reif Alison A²^ORCID

Affiliation:

1. Department of Statistics, Bioinformatics Research Center, North Carolina State University, Raleigh, NC 27695, USA

2. Biostatistics & Computational Biology Branch, National Institute of Environmental Health Sciences, Durham, NC 27709, USA

Abstract

Abstract Motivation The recently proposed knockoff filter is a general framework for controlling the false discovery rate (FDR) when performing variable selection. This powerful new approach generates a ‘knockoff’ of each variable tested for exact FDR control. Imitation variables that mimic the correlation structure found within the original variables serve as negative controls for statistical inference. Current applications of knockoff methods use linear regression models and conduct variable selection only for variables existing in model functions. Here, we extend the use of knockoffs for machine learning with boosted trees, which are successful and widely used in problems where no prior knowledge of model function is required. However, currently available importance scores in tree models are insufficient for variable selection with FDR control. Results We propose a novel strategy for conducting variable selection without prior model topology knowledge using the knockoff method with boosted tree models. We extend the current knockoff method to model-free variable selection through the use of tree-based models. Additionally, we propose and evaluate two new sampling methods for generating knockoffs, namely the sparse covariance and principal component knockoff methods. We test and compare these methods with the original knockoff method regarding their ability to control type I errors and power. In simulation tests, we compare the properties and performance of importance test statistics of tree models. The results include different combinations of knockoffs and importance test statistics. We consider scenarios that include main-effect, interaction, exponential and second-order models while assuming the true model structures are unknown. We apply our algorithm for tumor purity estimation and tumor classification using Cancer Genome Atlas (TCGA) gene expression data. Our results show improved discrimination between difficult-to-discriminate cancer types. Availability and implementation The proposed algorithm is included in the KOBT package, which is available at https://cran.r-project.org/web/packages/KOBT/index.html. Supplementary information Supplementary data are available at Bioinformatics online.

Funder

NIH

National Institute of Environmental Health Sciences

Publisher

Oxford University Press (OUP)

Subject

Computational Mathematics,Computational Theory and Mathematics,Computer Science Applications,Molecular Biology,Biochemistry,Statistics and Probability

Link

http://academic.oup.com/bioinformatics/advance-article-pdf/doi/10.1093/bioinformatics/btaa770/35334878/btaa770.pdf

Reference44 articles.

1. Systematic pan-cancer analysis of tumour purity;Aran;Nat. Commun,2015

2. Controlling the false discovery rate via knockoffs;Barber;Ann. Stat,2015

3. Hox genes and their role in the development of human cancers;Bhatlekar;J. Mol. Med,2014

4. Sparse estimation of a covariance matrix;Bien;Biometrika,2011

5. Random forests;Breiman;Mach. Learn,2001

Cited by 9 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. PLSKO: a robust knockoff generator to control false discovery rate in omics variable selection;2024-08-08

2. La replicabilidad en la ciencia y el papel transformador de la metodología estadística de knockoffs;SAHUARUS. REVISTA ELECTRÓNICA DE MATEMÁTICAS. ISSN: 2448-5365;2024-06-30

3. All that Glitters Is not Gold: Type‐I Error Controlled Variable Selection from Clinical Trial Data;Clinical Pharmacology & Therapeutics;2024-02-28

4. Epigenomic dissection of Alzheimer’s disease pinpoints causal variants and reveals epigenome erosion;Cell;2023-09

5. Human microglial state dynamics in Alzheimer’s disease progression;Cell;2023-09