In plain words: β-VAE pushes its internal code dimensions to be independent, but that is not the same as the hidden factors being independent. The analysis shows this mismatch makes results first improve then worsen as that push is turned up, so a middle setting works best.
Abstract · A Deeper Look at the Unsupervised Learning of Disentangled Representations in $β$-VAE from the Perspective of Core Object Recognition
The ability to recognize objects despite there being differences in appearance, known as Core Object Recognition, forms a critical part of human perception. While it is understood that the brain accomplishes Core Object Recognition through feedforward, hierarchical computations through the visual stream, the underlying algorithms that allow for invariant representations to form downstream is still not well understood. (DiCarlo et al., 2012) Various computational perceptual models have been built to attempt and tackle the object identification task in an artificial perceptual setting. Artificial Neural Networks, computational graphs consisting of weighted edges and mathematical operations at vertices, are loosely inspired by neural networks in the brain and have proven effective at various visual perceptual tasks, including object characterization and identification. (Pinto et al., 2008) (DiCarlo et al., 2012) For many data analysis tasks, learning representations where each dimension is statistically independent and thus disentangled from the others is useful. If the underlying generative factors of the data are also statistically independent, Bayesian inference of latent variables can form disentangled representations. This thesis constitutes a research project exploring a generalization of the Variational Autoencoder (VAE), $β$-VAE, that aims to learn disentangled representations using variational inference. $β$-VAE incorporates the hyperparameter $β$, and enforces conditional independence of its bottleneck neurons, which is in general not compatible with the statistical independence of latent variables. This text examines this architecture, and provides analytical and numerical arguments, with the goal of demonstrating that this incompatibility leads to a non-monotonic inference performance in $β$-VAE with a finite optimal $β$.
Harshvardhan Sikka
arXiv:2005.07114 · cs.LG, stat.ML · submitted Apr 25, 2020
abstract · pdf · 65 Pages, 6 Figures, Thesis
This thesis was the culmination of more than a year's worth of fulltime research. It was a blast to do, and in an exciting area. I've currently transitioned to a research scientist role at a defense company after completing my two Master's. I'd love to hear your thoughts or discuss anything!
I'm currently working on another DL project in my freetime with some collaborators. The goal is to build large sparse networks in a scalable way: https://www.harshsikka.com/creating-managing-and-understandi...