A modern maximum-likelihood theory for high-dimensional logistic regression-Reference-Cited by-同舟云学术

A modern maximum-likelihood theory for high-dimensional logistic regression

Published:2019-07-01 Issue:29 Volume:116 Page:14516-14525
ISSN:0027-8424
Container-title:Proceedings of the National Academy of Sciences
language:en
Short-container-title:Proc Natl Acad Sci USA

Author:

Sur Pragya,Candès Emmanuel J.

Abstract

Students in statistics or data science usually learn early on that when the sample size n is large relative to the number of variables p, fitting a logistic model by the method of maximum likelihood produces estimates that are consistent and that there are well-known formulas that quantify the variability of these estimates which are used for the purpose of statistical inference. We are often told that these calculations are approximately valid if we have 5 to 10 observations per unknown parameter. This paper shows that this is far from the case, and consequently, inferences produced by common software packages are often unreliable. Consider a logistic model with independent features in which n and p become increasingly large in a fixed ratio. We prove that (i) the maximum-likelihood estimate (MLE) is biased, (ii) the variability of the MLE is far greater than classically estimated, and (iii) the likelihood-ratio test (LRT) is not distributed as a χ2. The bias of the MLE yields wrong predictions for the probability of a case based on observed values of the covariates. We present a theory, which provides explicit expressions for the asymptotic bias and variance of the MLE and the asymptotic distribution of the LRT. We empirically demonstrate that these results are accurate in finite samples. Our results depend only on a single measure of signal strength, which leads to concrete proposals for obtaining accurate inference in finite samples through the estimate of this measure.

Funder

DOD | United States Navy | Office of Naval Research

National Science Foundation

Simons Foundation

Two Sigma

SU | School of Humanities and Sciences, Stanford University

Publisher

Proceedings of the National Academy of Sciences

Subject

Multidisciplinary

Reference39 articles.

1. The regression analysis of binary sequences;Cox;J. R. Stat. Soc. Ser. B (Methodol.),1958

2. D. W. Hosmer , S. Lemeshow , Applied Logistic Regression (John Wiley & Sons, 2013), vol. 398.

3. A. W. Van der Vaart , Asymptotic Statistics (Cambridge University Press), vol. 3.

4. R Core Team , R: A Language and Environment for Statistical Computing (R Foundation for Statistical Computing, Vienna, Austria, 2018).

5. The Large-Sample Distribution of the Likelihood Ratio for Testing Composite Hypotheses

Cited by 114 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Debiased lasso after sample splitting for estimation and inference in high‐dimensional generalized linear models;Canadian Journal of Statistics;2024-08-21

2. Determinants of local food producer participation in state‐sponsored marketing programs: Evidence from Missouri;Agribusiness;2024-08-13

3. A Non-Asymptotic Analysis of Generalized Vector Approximate Message Passing Algorithms With Rotationally Invariant Designs;IEEE Transactions on Information Theory;2024-08

4. On the functional regression model and its finite-dimensional approximations;Statistical Papers;2024-07-10

5. La replicabilidad en la ciencia y el papel transformador de la metodología estadística de knockoffs;SAHUARUS. REVISTA ELECTRÓNICA DE MATEMÁTICAS. ISSN: 2448-5365;2024-06-30