In plain words: Probability trees are branching diagrams of how causes lead to effects, and unlike standard causal networks they can show a cause matters only sometimes. New algorithms read them to answer what happened, what if we act, and what would have been, for any event.
Abstract
Probability trees are one of the simplest models of causal generative processes. They possess clean semantics and -- unlike causal Bayesian networks -- they can represent context-specific causal dependencies, which are necessary for e.g. causal induction. Yet, they have received little attention from the AI and ML community. Here we present concrete algorithms for causal reasoning in discrete probability trees that cover the entire causal hierarchy (association, intervention, and counterfactuals), and operate on arbitrary propositional and causal events. Our work expands the domain of causal reasoning to a very general class of discrete stochastic processes.
Tim Genewein, Tom McGrath, Grégoire Déletang, Vladimir Mikulik, Miljan Martic, Shane Legg, Pedro A. Ortega
arXiv:2010.12237 · cs.AI, cs.LG · submitted Oct 23, 2020 · updated Nov 12, 2020
abstract · pdf · html · (2nd version with correction to algorithm) 11 pages, 8 figures, 5 algorithms. A companion Colaboratory tutorial is available at https://github.com/deepmind/deepmind-research/tree/master/causal_reasoning
I'm quite happy to see more work on discrete generative models -- probabilistic programming languages are still wrestling (or simply ignoring!) the problem of "disintegration", where conditioning changes the base measure because it collapses the dimensionality of the probability manifold (similar to the issue Arjovsky identified with GANs). See e.g. http://homes.sice.indiana.edu/ccshan/rational/disint2arg.pdf and https://probprog.cc/assets/posters/thu/78.pdf. To me (although this is probably too radical a move to be palatable to most people) we should largely abandon continuous distributions and start building out discrete probability spaces and methods with the same vigor that continuous probability got in the form of measure theory. These probability trees some like a natural data structure to begin this. I'd also like to see representations for working with probabilities on discrete manifolds that approximate continuous space in computationally efficient ways -- there could be work in this direction already but I'm not aware of it.
Also, a fun implication of these probability trees I'd like to see explored is structural sharing: you need not copy an entire tree to represent the result of a conditioning or intervention. In general you need to copy only a number of nodes given by the size lying above the cut set, in a similar way to how immutable data structures like HAMTs can represent modified hash maps etc with persistent space efficiency by reusing unchanged nodes. If one expects to condition often with certain variables, it would then be useful to hoist such cut sets as high as possible -- does such a 'transpose' operation exist? I admit I did not read the paper thoroughly enough to know if this was mentioned.