In plain words: When many causes are studied at once, the method learns a hidden factor shared by them and uses it as a stand-in for unmeasured confounders. Needing weaker assumptions than the usual rule that all confounders are measured, it gave closer-to-truth effects in three studies.
Abstract
Causal inference from observational data often assumes "ignorability," that all confounders are observed. This assumption is standard yet untestable. However, many scientific studies involve multiple causes, different variables whose effects are simultaneously of interest. We propose the deconfounder, an algorithm that combines unsupervised machine learning and predictive model checking to perform causal inference in multiple-cause settings. The deconfounder infers a latent variable as a substitute for unobserved confounders and then uses that substitute to perform causal inference. We develop theory for the deconfounder, and show that it requires weaker assumptions than classical causal inference. We analyze its performance in three types of studies: semi-simulated data around smoking and lung cancer, semi-simulated data around genome-wide association studies, and a real dataset about actors and movie revenue. The deconfounder provides a checkable approach to estimating closer-to-truth causal effects.
Yixin Wang, David M. Blei
arXiv:1805.06826 · stat.ML, cs.LG, stat.ME · submitted May 17, 2018 · updated Apr 15, 2019
abstract · pdf · html · 72 pages