HIERARCHICAL SHRINKAGE IN TREE-BASED MODELS APPLIED TO IMAGE, HIGH-DIMENSIONAL, AND NOISY DATA

Loading...
Thumbnail Image

Date

Authors

Vazquez, Luis Eduardo

Journal Title

Journal ISSN

Volume Title

Publisher

University of Oklahoma – Graduate College

Abstract

Machine learning models often face a trade-off between predictive performance and interpretability—an issue of growing importance as AI systems become increasingly embedded in critical domains such as healthcare and finance. Tree-based models, like decision trees and random forests, are widely used due to their interpretability, but they are also known to overfit, especially in high-dimensional or imbalanced settings. This thesis investigates the use of \textit{hierarchical shrinkage}, a recently proposed post-hoc regularization technique, as a way to improve the generalization of decision tree-based models without compromising interpretability. The work of Agarwal et al.~\cite{DT:HS} focused on regression as well as on some tabular datasets for binary classification.%I begin by replicating the experiments from Agarwal et al.~\cite{DT:HS} to verify the behavior of hierarchical shrinkage before extending the method to new domains, which focused solely on tabular datasets. This motivated me to extended the evaluation of hierarchical shrinkage to new settings, not explored by Agarwal et al.~\cite{DT:HS}. First, I tested hierarchical shrinkage on image classification tasks. Our results indicate that, unlike in the tabular setting, hierarchical shrinkage does not improve performance on image data—regardless of whether dimensionality reduction via Principal Component Analysis (PCA) is applied, or not—and this holds for both decision trees and random forests. Next, I expanded the investigation to a broader range of tabular datasets, including both binary and multi-class problems. Our findings confirm that hierarchical shrinkage consistently improves performance on binary classification datasets and continues to offer improvements, though to a lesser extent, on multi-class datasets. I also observe that hierarchical shrinkage yields more substantial gains on high-dimensional tabular datasets. These trends hold across both individual decision trees and random forests. I further introduce synthetic label noise and apply confident learning to identify and denoise mislabeled instances. While confident learning leads to improved performance in some cases, especially when the noise is high, our experiments suggest that the impact of confident learning varies depending on the dataset and noise level. Moreover, some of the results that we obtain under noise are conflicting with the results that we obtain in the no-noise setting, thus indicating that more experimentation is needed here so that some safe conclusions can be drawn. Overall, this thesis provides a comprehensive evaluation of hierarchical shrinkage under a variety of learning conditions. Our results suggest that hierarchical shrinkage is a promising, interpretable regularization method for decision tree-based models.

Description

Citation

Related file

Notes

Collections

Endorsement

Review

Supplemented By

Referenced By

DOI

Collection Detail

# of Isolates from RBM

# of Isolates from TV8