In plain words: A system that restyles images without paired examples or even labels saying which style each image belongs to, by figuring out the image groups and the translations at the same time. It matched or beat a version trained with full labels and stayed reliable even when the guessed number of groups was off.
Abstract
Every recent image-to-image translation model inherently requires either image-level (i.e. input-output pairs) or set-level (i.e. domain labels) supervision. However, even set-level supervision can be a severe bottleneck for data collection in practice. In this paper, we tackle image-to-image translation in a fully unsupervised setting, i.e., neither paired images nor domain labels. To this end, we propose a truly unsupervised image-to-image translation model (TUNIT) that simultaneously learns to separate image domains and translates input images into the estimated domains. Experimental results show that our model achieves comparable or even better performance than the set-level supervised model trained with full labels, generalizes well on various datasets, and is robust against the choice of hyperparameters (e.g. the preset number of pseudo domains). Furthermore, TUNIT can be easily extended to semi-supervised learning with a few labeled data.
Kyungjune Baek, Yunjey Choi, Youngjung Uh, Jaejun Yoo, Hyunjung Shim
arXiv:2006.06500 · cs.CV, cs.LG · submitted Jun 11, 2020 · updated Aug 20, 2021
abstract · pdf · html · Accepted to ICCV 2021