On clustering levels of a hierarchical categorical risk factor-Reference-Cited by-同舟云学术

On clustering levels of a hierarchical categorical risk factor

Published:2024-02-01 Issue: Volume: Page:1-39
ISSN:1748-4995
Container-title:Annals of Actuarial Science
language:en
Short-container-title:Ann. actuar. sci.

Author:

Campo Bavo D.C.^ORCID,Antonio Katrien

Abstract

Abstract Handling nominal covariates with a large number of categories is challenging for both statistical and machine learning techniques. This problem is further exacerbated when the nominal variable has a hierarchical structure. We commonly rely on methods such as the random effects approach to incorporate these covariates in a predictive model. Nonetheless, in certain situations, even the random effects approach may encounter estimation problems. We propose the data-driven Partitioning Hierarchical Risk-factors Adaptive Top-down algorithm to reduce the hierarchically structured risk factor to its essence, by grouping similar categories at each level of the hierarchy. We work top-down and engineer several features to characterize the profile of the categories at a specific level in the hierarchy. In our workers’ compensation case study, we characterize the risk profile of an industry via its observed damage rates and claim frequencies. In addition, we use embeddings to encode the textual description of the economic activity of the insured company. These features are then used as input in a clustering algorithm to group similar categories. Our method substantially reduces the number of categories and results in a grouping that is generalizable to out-of-sample data. Moreover, we obtain a better differentiation between high-risk and low-risk companies.

Publisher

Cambridge University Press (CUP)

Subject

Statistics, Probability and Uncertainty,Economics and Econometrics,Statistics and Probability

Reference104 articles.

1. A note on a hierarchical interpretation for negative variance components;Molenberghs;Statistical Modelling,2011

2. Rentzmann, S. & Wuthrich, M. V. (2019). Unsupervised learning: what is a sports car? Available at: https://ssrn.com/abstract=3439358 or 10.2139/ssrn.3439358.

3. Using clusters based on social determinants to identify the top 5% utilizers of health care;Rosenberg;North American Actuarial Journal,2022

4. Approximate inference in generalized linear mixed models;Breslow;Journal of the American Statistical Association,1993

5. Clustering in an object-oriented environment;Struyf;Journal of Statistical Software,1997