Using Interpretable Machine Learning for Differential Item Functioning Detection in Psychometric Tests-Reference-Cited by-同舟云学术

Using Interpretable Machine Learning for Differential Item Functioning Detection in Psychometric Tests

Published:2024-03-11 Issue:4-5 Volume:48 Page:167-186
ISSN:0146-6216
Container-title:Applied Psychological Measurement
language:en
Short-container-title:Applied Psychological Measurement

Author:

Kraus Elisabeth Barbara¹^ORCID,Wild Johannes²,Hilbert Sven²

Affiliation:

1. LMU Munich, Germany

2. University of Regensburg, Germany

Abstract

This study presents a novel method to investigate test fairness and differential item functioning combining psychometrics and machine learning. Test unfairness manifests itself in systematic and demographically imbalanced influences of confounding constructs on residual variances in psychometric modeling. Our method aims to account for resulting complex relationships between response patterns and demographic attributes. Specifically, it measures the importance of individual test items, and latent ability scores in comparison to a random baseline variable when predicting demographic characteristics. We conducted a simulation study to examine the functionality of our method under various conditions such as linear and complex impact, unfairness and varying number of factors, unfair items, and varying test length. We found that our method detects unfair items as reliably as Mantel–Haenszel statistics or logistic regression analyses but generalizes to multidimensional scales in a straight forward manner. To apply the method, we used random forests to predict migration backgrounds from ability scores and single items of an elementary school reading comprehension test. One item was found to be unfair according to all proposed decision criteria. Further analysis of the item’s content provided plausible explanations for this finding. Analysis code is available at: https://osf.io/s57rw/?view_only=47a3564028d64758982730c6d9c6c547 .

Funder

Bayerisches Staatsministerium für Bildung und Kultus, Wissenschaft und Kunst

Publisher

SAGE Publications

Link

https://journals.sagepub.com/doi/pdf/10.1177/01466216241238744

Reference52 articles.

1. Guess Where: The Position of Correct Answers in Multiple-Choice Test Items as a Psychometric Variable

2. Simplifying the Assessment of Measurement Invariance over Multiple Background Variables: Using Regularized Moderated Nonlinear Factor Analysis to Detect Differential Item Functioning

3. The Development of Cognitive, Language, and Cultural Skills From Age 3 to 6

4. Improving the assessment of measurement invariance: Using regularization to select anchor items and identify differential item functioning.

5. The Multidimensionality of Measurement Bias in High‐Stakes Testing: Using Machine Learning to Evaluate Complex Sources of Differential Item Functioning