about
Video2StyleGAN: Disentangling Local and Global Variations in a Video (arxiv.org)
50 points by lnyan on May 30, 2022 | hide | past | pdf | 5 comments on HN

In plain words: A system copies a person's head position, pose, and facial expressions from a driving video onto a target face by splitting those controls across separate internal settings of a pretrained face generator and combining them. It beat other approaches in tough scenarios.

Abstract

Image editing using a pretrained StyleGAN generator has emerged as a powerful paradigm for facial editing, providing disentangled controls over age, expression, illumination, etc. However, the approach cannot be directly adopted for video manipulations. We hypothesize that the main missing ingredient is the lack of fine-grained and disentangled control over face location, face pose, and local facial expressions. In this work, we demonstrate that such a fine-grained control is indeed achievable using pretrained StyleGAN by working across multiple (latent) spaces (namely, the positional space, the W+ space, and the S space) and combining the optimization results across the multiple spaces. Building on this enabling component, we introduce Video2StyleGAN that takes a target image and driving video(s) to reenact the local and global locations and expressions from the driving video in the identity of the target image. We evaluate the effectiveness of our method over multiple challenging scenarios and demonstrate clear improvements over alternative approaches.

Rameen Abdal, Peihao Zhu, Niloy J. Mitra, Peter Wonka
arXiv:2205.13996 · cs.CV, cs.GR · submitted May 27, 2022 · updated May 30, 2022
abstract · pdf · html · Video : https://youtu.be/oUeXFyfdE1A

add comment on HN

AI image manipulation and creation is getting really wild. The last 12 months has seen a real explosion in published research. It feels like an exponential curve.
Absolutely. I feel like I never have a good understanding of what is possible anymore because every few months something shows up that blows away all my expectations of what tech can do.

It feels like not so long ago the best we had was the google deep dream stuff that was only good at generating weird nightmare images of dog faces.

Is there anything more useless than these single digit page pdfs full of irreproducible claims?
The paper seems to be about generating videos, where can I view those videos?
"Comments: Video: this https URL" -> https://youtu.be/oUeXFyfdE1A