Abstract
In this dissertation, we study three analytical goals related to imbalanced data: (1) determining the population class distribution when it is difficult to obtain sufficient observations on one of the classes; (2) comparing explanatory and predictive modeling of imbalanced data using logistic regression, discriminant analysis, and decision trees; and (3) considering suitable performance evaluation metrics for different purposes. Our research question is focused on comparing weighting and intercept correction for undersampled data using logistic regression, discriminant analysis, and decision trees (three models) when dealing with imbalanced data. Specifically, we study the following: (1) The relationship between sampling data with a weighted decision tree and an unweighted decision tree. (2) What are rules to induce a correction of the probability cutoff in a decision tree for obtaining the population model? (3) What is the difference between explanatory and predictive modeling? We study these questions using several datasets with different corrections. We find that when building training models across different data distributions, when the imbalanced data set is very large, using weighting and intercept correction with logistic regression and discriminant analysis lead to consistent results in both explanatory and predictive tasks. In contrast, decision trees display different results when we investigate explanatory factors (tree variables) and predicting or ranking new observations. Therefore, we should carefully deal with imbalanced data using decision trees.