In plain words: A mathematical test predicts when a little extra help—sparse labels, matched pairs, or ranked pairs—makes a system separate the hidden factors behind the data. Experiments confirmed its predictions about which kinds of help guarantee that separation and which do not.
Abstract
Learning disentangled representations that correspond to factors of variation in real-world data is critical to interpretable and human-controllable machine learning. Recently, concerns about the viability of learning disentangled representations in a purely unsupervised manner has spurred a shift toward the incorporation of weak supervision. However, there is currently no formalism that identifies when and how weak supervision will guarantee disentanglement. To address this issue, we provide a theoretical framework to assist in analyzing the disentanglement guarantees (or lack thereof) conferred by weak supervision when coupled with learning algorithms based on distribution matching. We empirically verify the guarantees and limitations of several weak supervision methods (restricted labeling, match-pairing, and rank-pairing), demonstrating the predictive power and usefulness of our theoretical framework.
Rui Shu, Yining Chen, Abhishek Kumar, Stefano Ermon, Ben Poole
arXiv:1910.09772 · cs.LG, stat.ML · submitted Oct 22, 2019 · updated Apr 10, 2020
abstract · pdf · html · ICLR 2020