In plain words: It trains one model across many datasets, forcing the same classifier to work best in each, keeping only patterns that hold everywhere. Theory and tests show these patterns track real causes and carry to new data, unlike ordinary training that fits the training set.
Abstract · Invariant Risk Minimization
We introduce Invariant Risk Minimization (IRM), a learning paradigm to estimate invariant correlations across multiple training distributions. To achieve this goal, IRM learns a data representation such that the optimal classifier, on top of that data representation, matches for all training distributions. Through theory and experiments, we show how the invariances learned by IRM relate to the causal structures governing the data and enable out-of-distribution generalization.
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, David Lopez-Paz
arXiv:1907.02893 · stat.ML, cs.AI, cs.LG · submitted Jul 5, 2019 · updated Mar 27, 2020
abstract · pdf · html