about
Which Training Methods for GANs Do Converge? (2018) (arxiv.org)
1 point by QuadmasterXLII on Mar 24, 2022 | hide | past | pdf | 1 comment on HN

In plain words: Generative adversarial networks pit a generator against a checker to make images; the study tests which training tricks settle on the answer when data sits on thin, lower-dimensional surfaces. Plain training and the Wasserstein approach can fail, while adding noise or gradient penalties converges.

Abstract · Which Training Methods for GANs do actually Converge?

Recent work has shown local convergence of GAN training for absolutely continuous data and generator distributions. In this paper, we show that the requirement of absolute continuity is necessary: we describe a simple yet prototypical counterexample showing that in the more realistic case of distributions that are not absolutely continuous, unregularized GAN training is not always convergent. Furthermore, we discuss regularization strategies that were recently proposed to stabilize GAN training. Our analysis shows that GAN training with instance noise or zero-centered gradient penalties converges. On the other hand, we show that Wasserstein-GANs and WGAN-GP with a finite number of discriminator updates per generator update do not always converge to the equilibrium point. We discuss these results, leading us to a new explanation for the stability problems of GAN training. Based on our analysis, we extend our convergence results to more general GANs and prove local convergence for simplified gradient penalties even if the generator and data distribution lie on lower dimensional manifolds. We find these penalties to work well in practice and use them to learn high-resolution generative image models for a variety of datasets with little hyperparameter tuning.

Lars Mescheder, Andreas Geiger, Sebastian Nowozin
arXiv:1801.04406 · cs.LG, cs.AI, cs.GT · submitted Jan 13, 2018 · updated Jul 31, 2018
abstract · pdf · html · conference

add comment on HN

In the past, when I read Generative Adversarial Network research, I found a veritable zoo of approaches to regularizing the discriminator, such as Least-Squares-GAN, Wasserstein-GAN, Wasserstein-Gradient-Penalty, and Instance Noise. However, in the past few years I’ve noticed that the simple gradient penalty labelled ‘R1’ suddenly became so dominant that papers using it don’t draw attention to it, they just say ‘R1 regularization’ and cite this work.