about
Audio-Driven Emotional 3D Talking-Head Generation (arxiv.org)
2 points by sandwichsphinx on Oct 24, 2024 | hide | past | pdf | discuss on HN

In plain words: It turns speech into a map of facial landmarks, blends in an emotion signal to shape the face, then renders those landmarks into a lifelike video portrait. Unlike the usual focus on lip sync and image sharpness alone, it produces more accurate emotional expressions than earlier systems.

Abstract · EmoGene: Audio-Driven Emotional 3D Talking-Head Generation

Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelity and lip synchronization, generating accurate emotional expressions remains underexplored. In this paper, we introduce EmoGene, a novel framework for synthesizing high-fidelity, audio-driven video portraits with accurate emotional expressions. Our approach employs a variational autoencoder (VAE)-based audio-to-motion module to generate facial landmarks, which are concatenated with emotional embedding in a motion-to-emotion module to produce emotional landmarks. These landmarks drive a Neural Radiance Fields (NeRF)-based emotion-to-video module to render realistic emotional talking-head videos. Additionally, we propose a pose sampling method to generate natural idle-state (non-speaking) videos for silent audio inputs. Extensive experiments demonstrate that EmoGene outperforms previous methods in generating high-fidelity emotional talking-head videos.

Wenqing Wang, Yun Fu
arXiv:2410.17262 · cs.CV, cs.AI, cs.HC, cs.LG · submitted Oct 7, 2024 · updated May 1, 2025
abstract · pdf · html · Accepted by the 2025 IEEE 19th International Conference on Automatic Face and Gesture Recognition (FG)

add comment on HN