DIAGNOSTIC CLASSIFICATION IN THE PRESENCE OF INTERSECTIONAL DIFFERENTIAL ITEM FUNCTIONING: A MONTE CARLO STUDY ON IRT AND RANDOM FOREST APPROACHES
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Psychological researchers frequently use self-report scales to classify individuals into diagnostic categories. Item response theory (IRT) is the recommended family of mathematical models for scoring psychological scales. The accuracy of IRT scores is dependent upon the extent to which the assumptions invoked by the IRT model are met. Techniques such as multi-group IRT allow for more accurate estimation of scores when the assumption of item invariance is violated. However, if the grouping variable causing differential item functioning (DIF) is complex, unmeasured, or unknown, modeling DIF through the IRT framework becomes unreasonable. Random forests (RF), as a nonparametric machine learning approach, do not assume item invariance and may be a reasonable approach to classification when DIF is complex, unmeasured, or unknown. Using a Monte Carlo simulation, this study examined the performance of three IRT-based approaches and one RF-based approach to classification when data contained complex DIF caused by multiple demographic variables. Classification performance was evaluated across six metrics: accuracy, sensitivity, specificity, precision, negative predictive value, and F1 score. Results indicate that the classification performance of the second-order IRT model-based approach was higher than that of either of the multidimensional IRT model-based approaches. Higher classification performance was observed from the RF-based approach than the second-order IRT-based approach in most conditions. However, in conditions where the focal group was simulated to have the same or higher levels of the latent trait than the reference group, second-order IRT-based classifications were more sensitive than RF classifications. Additionally, the classification performance of the RF approach was strongly associated with measurement error, such that more measurement error led to lower classification performance.