In plain words: A decision tree gives missing values a third branch, assuming they say nothing about the answer. It beats sending missing values down whichever branch helps most when values go missing at random, especially in new data, but not when being missing is a clue.
Abstract · Trinary Decision Trees for handling missing data
This paper introduces the Trinary decision tree, an algorithm designed to improve the handling of missing data in decision tree regressors and classifiers. Unlike other approaches, the Trinary decision tree does not assume that missing values contain any information about the response. Both theoretical calculations on estimator bias and numerical illustrations using real data sets are presented to compare its performance with established algorithms in different missing data scenarios (Missing Completely at Random (MCAR), and Informative Missingness (IM)). Notably, the Trinary tree outperforms its peers in MCAR settings, especially when data is only missing out-of-sample, while lacking behind in IM settings. A hybrid model, the TrinaryMIA tree, which combines the Trinary tree and the Missing In Attributes (MIA) approach, shows robust performance in all types of missingness. Despite the potential drawback of slower training speed, the Trinary tree offers a promising and more accurate method of handling missing data in decision tree algorithms.
Henning Zakrisson
arXiv:2309.03561 · stat.ML, cs.LG · submitted Sep 7, 2023 · updated Jan 11, 2024
abstract · pdf · html
> "Notably, the Trinary tree outperforms its peers in MCAR settings, especially when data is only missing out-of-sample, while lacking behind in IM settings."
This somewhat mirrors the behavior of early imputation strategies. One must ponder, however, how the Trinary tree would perform vis-a-vis older methods like CART's surrogate splits or C4.5's probabilistic splits for handling missing values. These older methods were crafted with an intuition somewhat similar to the Trinary tree.
It's also great to see the amalgamation of Trinary tree with the Missing In Attributes approach into the TrinaryMIA tree. But the efficacy of this hybrid model isn't completely surprising. MIA has historically shown resilience in diverse missing data scenarios, and combining that with the Trinary's approach could harmonize their strengths.
What would be really enticing is to see if the essence of the Trinary decision tree can be injected into boosting models like XGBoost or LightGBM. Since these models are notorious for their treatment of missing values, maybe there's some potential symbiosis there?