Affiliation:
1. Institute of Statistics, Biostatistics and Actuarial Sciences LIDAM Louvain‐la‐Neuve Belgium
2. Department of Mathematics, Computer Science, and Natural Sciences University of Hamburg Hamburg Germany
Abstract
Sparse linear prediction methods suffer from decreased prediction accuracy when the predictor variables have cluster structure (e.g., highly correlated groups of variables). To improve prediction accuracy, various methods have been proposed to identify variable clusters from the data and integrate cluster information into a sparse modeling process. But none of these methods achieve satisfactory performance for prediction, variable selection and variable clustering performed simultaneously. This paper presents Variable Cluster Principal Component Regression (VC‐PCR), a prediction method that uses variable selection and variable clustering in order to solve this problem. Experiments with real and simulated data demonstrate that, compared to competitor methods, VC‐PCR is the only method that achieves simultaneously good prediction, variable selection, and clustering performance when cluster structure is present.