about
NPGA: Neural Parametric Gaussian Avatars – high-fidelity digital faces (arxiv.org)
60 points by samspenc on May 30, 2024 | hide | past | pdf | 14 comments on HN

In plain words: Builds lifelike, animatable head avatars from multi-view video by driving tiny translucent 3D dots with a head model that captures many expressions, then learning extra wrinkles from the footage. It beat the best earlier avatars at re-animating someone by 2.6 points on a photo-quality score.

Abstract · NPGA: Neural Parametric Gaussian Avatars

The creation of high-fidelity, digital versions of human heads is an important stepping stone in the process of further integrating virtual components into our everyday lives. Constructing such avatars is a challenging research problem, due to a high demand for photo-realism and real-time rendering performance. In this work, we propose Neural Parametric Gaussian Avatars (NPGA), a data-driven approach to create high-fidelity, controllable avatars from multi-view video recordings. We build our method around 3D Gaussian splatting for its highly efficient rendering and to inherit the topological flexibility of point clouds. In contrast to previous work, we condition our avatars' dynamics on the rich expression space of neural parametric head models (NPHM), instead of mesh-based 3DMMs. To this end, we distill the backward deformation field of our underlying NPHM into forward deformations which are compatible with rasterization-based rendering. All remaining fine-scale, expression-dependent details are learned from the multi-view videos. For increased representational capacity of our avatars, we propose per-Gaussian latent features that condition each primitives dynamic behavior. To regularize this increased dynamic expressivity, we propose Laplacian terms on the latent features and predicted dynamics. We evaluate our method on the public NeRSemble dataset, demonstrating that NPGA significantly outperforms the previous state-of-the-art avatars on the self-reenactment task by 2.6 PSNR. Furthermore, we demonstrate accurate animation capabilities from real-world monocular videos.

Simon Giebenhain, Tobias Kirschstein, Martin Rünz, Lourdes Agapito, Matthias Nießner
arXiv:2405.19331 · cs.CV, cs.AI, cs.GR · submitted May 29, 2024 · updated Sep 13, 2024
abstract · pdf · html · Project Page: see https://simongiebenhain.github.io/NPGA/ ; Youtube Video: see https://youtu.be/t0S0OK7WnA4

add comment on HN

Their github page has some videos: https://simongiebenhain.github.io/NPGA/
Um, wow. These are really, really good. They are not perfect, but the improvements on fidelity, open mouth, eyes over a GAN-based approach are .. real high.

This is the first paper I’ve seen with videos that compare a person with a re-render side by side, and it’s a nice way to see what the model’s good at, and what it’s not.

Some perf numbers (which they say are unoptimized): 30-60 hrs on a 3080 for the avatar model, and rendering in the 20-40fps range on the same hardware. Basically good enough for a commercial implementation. They don’t mention latency of the CNN side that I can find, which is obviously a big question for chat scenarios, although not a big deal for pre-rendered scenes.

Agreed, very impressive results. It's both ~worrying and amazing that I'm sure an AI agent could just directly trace a path in the expression latent space to produce a photorealistic and real-time rendered head.
it feels to me like the gaussian-based algorithm improvements will likely force new render pipelines away from triangles sooner rather than later. It's hard to imagine giving up fidelity like this. And re-rendering to textured triangles is not fast right now. Should be a fun couple of years!
The real issue is the dependency in a real life enactment. Not even considering the performance.
In the novel Snow Crash, one of the characters, Juanita, does just this - she is known for making realistic facial expressions possible in the metaverse.
Given the now clear abuse potential (and other issues surrounding AI generally), why do papers like this discuss the motivation as if it’s a for sure requirement to do such work?

“The creation of high-fidelity, digital versions of human heads is an important stepping stone in the process of further integrating virtual components into our everyday lives.”

It’s not clear to me that’s a desirable outcome.

It is kind of a requirement. Scientific papers will typically describe the motivation for a work to start the introduction. The real motivation of course is “I need to publish something”.
The real motivation is "I want to research this field", publishing is part of the way. Nobody is doing research "to publish something", there are much simpler fields to enter if that was the motivation.
If you decide to research in $field, then you’ll typically need to keep publishing papers to keep your job
Indeed. That's what I'm saying - you decide you want to research and so you publish. Nobody decides to publish and so does research.
How is this a useful critique? The publishing being a subgoal doesn’t mean it’s not still a goal.
How is that a useful critique? Yes, researchers publish research.
I don't understand your question. There are some cases of abuse and so the researchers shouldn't discuss applications of the research?