In plain words: Decision trees and diffusion models are the same process in the right limits, sharing one training rule that gradient boosting follows almost perfectly. A generator built on this idea makes tabular data as faithfully as the usual approach but twice as fast.
Abstract
Decision trees and diffusion models are ostensibly disparate model classes, one discrete and hierarchical, the other continuous and dynamic. This work unifies the two by establishing a crisp mathematical correspondence between hierarchical decision trees and diffusion processes in appropriate limiting regimes. Our unification reveals a shared optimization principle: \emph{Global Trajectory Score Matching (GTSM)}, for which gradient boosting (in an idealized version) is asymptotically optimal. We underscore the conceptual value of our work through two key practical instantiations: \treeflow, which achieves competitive generation quality on tabular data with higher fidelity and a 2\times computational speedup, and \dsmtree, a novel distillation method that transfers hierarchical decision logic into neural networks, matching teacher performance within 2\% on many benchmarks.
Sai Niranjan Ramachandran, Suvrit Sra
arXiv:2605.00414 · cs.LG, cond-mat.stat-mech, cs.AI · submitted May 1, 2026 · updated May 21, 2026
abstract · pdf · html · 12 pages (main), 68 pages (inclusive of appendix), Accepted in the Forty-Third International Conference on Machine Learning (ICML) 2026
Do we think they'll be better than decision trees? Is there some tabular problem that can be handled by diffusion but not trees?