about
Facial Performance Capture with Deep Neural Networks (arxiv.org)
3 points by metafunctor on Sep 23, 2016 | hide | past | pdf | 1 comment on HN

In plain words: Film an actor with a multi-camera rig for 5-10 minutes and a neural network learns their face well enough to rebuild its full 3D shape and motion from single-camera video, including hidden parts. It tracked eyes and lips more convincingly than other real-time single-camera trackers.

Abstract · Production-Level Facial Performance Capture Using Deep Convolutional Neural Networks

We present a real-time deep learning framework for video-based facial performance capture -- the dense 3D tracking of an actor's face given a monocular video. Our pipeline begins with accurately capturing a subject using a high-end production facial capture pipeline based on multi-view stereo tracking and artist-enhanced animations. With 5-10 minutes of captured footage, we train a convolutional neural network to produce high-quality output, including self-occluded regions, from a monocular video sequence of that subject. Since this 3D facial performance capture is fully automated, our system can drastically reduce the amount of labor involved in the development of modern narrative-driven video games or films involving realistic digital doubles of actors and potentially hours of animated dialogue per character. We compare our results with several state-of-the-art monocular real-time facial capture techniques and demonstrate compelling animation inference in challenging areas such as eyes and lips.

Samuli Laine, Tero Karras, Timo Aila, Antti Herva, Shunsuke Saito, Ronald Yu, Hao Li, Jaakko Lehtinen
arXiv:1609.06536 · cs.CV, cs.GR · submitted Sep 21, 2016 · updated Jun 2, 2017
abstract · pdf · html · Final SCA 2017 version

add comment on HN

In other words: transfer of video footage into a motion sequence of a 3D mesh representing an actor's face.