A simulation study of the performance of statistical models for count outcomes with excessive zeros-Reference-Cited by-同舟云学术

A simulation study of the performance of statistical models for count outcomes with excessive zeros

Published:2024-08-28 Issue: Volume: Page:
ISSN:0277-6715
Container-title:Statistics in Medicine
language:en
Short-container-title:Statistics in Medicine

Author:

Zhou Zhengyang¹^ORCID,Li Dateng²^ORCID,Huh David³,Xie Minge⁴,Mun Eun‐Young¹

Affiliation:

1. Department of Population and Community Health University of North Texas Health Science Center Fort Worth Texas USA

2. Norden Lofts White Plains New York USA

3. School of Social Work University of Washington Seattle Washington USA

4. Department of Statistics Rutgers University Piscataway New Jersey USA

Abstract

Background: Outcome measures that are count variables with excessive zeros are common in health behaviors research. Examples include the number of standard drinks consumed or alcohol‐related problems experienced over time. There is a lack of empirical data about the relative performance of prevailing statistical models for assessing the efficacy of interventions when outcomes are zero‐inflated, particularly compared with recently developed marginalized count regression approaches for such data. Methods: The current simulation study examined five commonly used approaches for analyzing count outcomes, including two linear models (with outcomes on raw and log‐transformed scales, respectively) and three prevailing count distribution‐based models (ie, Poisson, negative binomial, and zero‐inflated Poisson (ZIP) models). We also considered the marginalized zero‐inflated Poisson (MZIP) model, a novel alternative that estimates the overall effects on the population mean while adjusting for zero‐inflation. Motivated by alcohol misuse prevention trials, extensive simulations were conducted to evaluate and compare the statistical power and Type I error rate of the statistical models and approaches across data conditions that varied in sample size ( to 500), zero rate (0.2 to 0.8), and intervention effect sizes. Results: Under zero‐inflation, the Poisson model failed to control the Type I error rate, resulting in higher than expected false positive results. When the intervention effects on the zero (vs. non‐zero) and count parts were in the same direction, the MZIP model had the highest statistical power, followed by the linear model with outcomes on the raw scale, negative binomial model, and ZIP model. The performance of the linear model with a log‐transformed outcome variable was unsatisfactory. Conclusions: The MZIP model demonstrated better statistical properties in detecting true intervention effects and controlling false positive results for zero‐inflated count outcomes. This MZIP model may serve as an appealing analytical approach to evaluating overall intervention effects in studies with count outcomes marked by excessive zeros.

Funder

National Science Foundation of Sri Lanka

National Institute on Alcohol Abuse and Alcoholism

National Institute of Allergy and Infectious Diseases

Publisher

Wiley

Link

https://onlinelibrary.wiley.com/doi/pdf/10.1002/sim.10198

Reference23 articles.

1. A tutorial on individual participant data meta-analysis using Bayesian multilevel modeling to estimate alcohol intervention effects across heterogeneous studies

2. The effect of a major cigarette price change on smoking behavior in california: a zero-inflated negative binomial model

3. The role of mother–daughter sexual risk communication in reducing sexual risk behaviors among urban adolescent females: a prospective study

4. Brief Motivational Interventions for College Student Drinking May Not Be as Powerful as We Think: An Individual Participant-Level Data Meta-Analysis

5. A bias correction method in meta‐analysis of randomized clinical trials with no adjustments for zero‐inflated outcomes