about
Estimating the probabilities of causation via deep monotonic twin networks (arxiv.org)
40 points by Anon84 on Sep 11, 2021 | hide | past | pdf | 5 comments on HN

In plain words: They set a rule for how cause-and-effect links should behave when outcomes have several categories, then train neural networks following it to answer "what would have happened if" questions. This gave accurate estimates on medical, disease and finance data; without the rule, results were wrong.

Abstract · Estimating Categorical Counterfactuals via Deep Twin Networks

Counterfactual inference is a powerful tool, capable of solving challenging problems in high-profile sectors. To perform counterfactual inference, one requires knowledge of the underlying causal mechanisms. However, causal mechanisms cannot be uniquely determined from observations and interventions alone. This raises the question of how to choose the causal mechanisms so that resulting counterfactual inference is trustworthy in a given domain. This question has been addressed in causal models with binary variables, but the case of categorical variables remains unanswered. We address this challenge by introducing for causal models with categorical variables the notion of counterfactual ordering, a principle that posits desirable properties causal mechanisms should posses, and prove that it is equivalent to specific functional constraints on the causal mechanisms. To learn causal mechanisms satisfying these constraints, and perform counterfactual inference with them, we introduce deep twin networks. These are deep neural networks that, when trained, are capable of twin network counterfactual inference -- an alternative to the abduction, action, & prediction method. We empirically test our approach on diverse real-world and semi-synthetic data from medicine, epidemiology, and finance, reporting accurate estimation of counterfactual probabilities while demonstrating the issues that arise with counterfactual reasoning when counterfactual ordering is not enforced.

Athanasios Vlontzos, Bernhard Kainz, Ciaran M. Gilligan-Lee
arXiv:2109.01904 · cs.LG, cs.AI · submitted Sep 4, 2021 · updated Jan 20, 2023
abstract · pdf · html

add comment on HN

TL;DR: this looks like it is about methods to answer questions like "Was it event X that caused event Y?"

This feels difficult for a layman like me. Let's try to clear that up.

> However, as noted by Pearl, interventional queries only form part of a larger hierarchy of causal queries, with counterfactuals sitting at the top.

Regarding "Pearl", paper text cites:

> Tian, J.; and Pearl, J. 2000. Probabilities of causation: Bounds and identification. Annals of Mathematics and Artificial Intelligence, 28(1): 287–313.

which appears to be this: https://link.springer.com/article/10.1023/A:1018912507879

and appears to provide mathematical foundations to answering questions like "did event A cause event B" from observations.

Regarding the "hierarchy", this looks like a clear and very short introduction: http://web.cs.ucla.edu/~kaoru/3-layer-causal-hierarchy.pdf

From the very end of the paper:

> Causal inference is a tool that can have significant impact on society depending on its use. As such the authors are adamant that all uses of causal inference that could have negative societal impacts should be accompanied with the proper due diligence and fail-safes in order to minimize and even eliminate said negative impacts.

Does this translate to "Applications of this research might destroy society, please refrain"?

>Does this translate to "Applications of this research might destroy society, please refrain"?

The biggest potential use of casual inference from what I remember in school is social policy predictions. Large scale changes that you simply cannot randomly test. Things like "does prosecuting minor crimes more heavily lower overall crime" or "does privatizing schools increase student test scores." Even if you get the causality right in the main metric there may be secondary effects you don't even test for.

The main point about most causality questions is to answer the question “what if X did not happen / did happen”.

Normally you can’t get that information since you are looking at past events where X happened and you can’t change that. In addition to that it also helps answering questions where you could simulate a different behavior (as in a controlled random trial), but it’s simply too expensive / unethical / … (Is smoking really causing cancer, or not —- would X have had cancer if they didn’t smoke). Answering these kind of questions in a mathematical way gives a lot of power since it removes / kind of solves the correlation doesn’t equal causation problem

One application I can think of would be to dialog systems. Tracing the meaning in a conversation between two strangers is determined by statements and responses, and the degree of unexpectedness of each utterance. Say a person asks a yes or no question, but gets a response along the lines of "I'm talking about X, not Y"; it would be useful to be able to go back and retrace the conversation at that point and try different responses in an effort understand what's being said.
> Does this translate to "Applications of this research might destroy society, please refrain"?

This is boilerplate acknowledgement of ethics concerns that often appears in DL/ML papers these days. Doesn't really mean anything.