about
Yoshua Bengio: Learning Causal Models Online (arxiv.org)
1 point by Anon84 on Jun 16, 2020 | hide | past | pdf | discuss on HN

In plain words: A learning system watches how each feature's link to the answer changes over time and drops the ones that keep shifting, since those are unreliable shortcuts. This keeps only steady features and generalizes well, but only if the data keeps its original time order.

Abstract · Learning Causal Models Online

Predictive models -- learned from observational data not covering the complete data distribution -- can rely on spurious correlations in the data for making predictions. These correlations make the models brittle and hinder generalization. One solution for achieving strong generalization is to incorporate causal structures in the models; such structures constrain learning by ignoring correlations that contradict them. However, learning these structures is a hard problem in itself. Moreover, it's not clear how to incorporate the machinery of causality with online continual learning. In this work, we take an indirect approach to discovering causal models. Instead of searching for the true causal model directly, we propose an online algorithm that continually detects and removes spurious features. Our algorithm works on the idea that the correlation of a spurious feature with a target is not constant over-time. As a result, the weight associated with that feature is constantly changing. We show that by continually removing such features, our method converges to solutions that have strong generalization. Moreover, our method combined with random search can also discover non-spurious features from raw sensory data. Finally, our work highlights that the information present in the temporal structure of the problem -- destroyed by shuffling the data -- is essential for detecting spurious features online.

Khurram Javed, Martha White, Yoshua Bengio
arXiv:2006.07461 · cs.LG, cs.AI, stat.ML · submitted Jun 12, 2020
abstract · pdf · html · Spurious features, causal models, online learning, random search, non-iid

add comment on HN