about
WonderHuman: 3D avatars from single-view video (arxiv.org)
38 points by jinqueeny on Feb 20, 2025 | hide | past | pdf | 6 comments on HN

In plain words: It builds a moving 3D human avatar from a one-camera video, using a picture generator to invent body parts the camera never saw while keeping them consistent with the footage. It renders unseen parts more photorealistically than methods that need the whole body filmed.

Abstract · WonderHuman: Hallucinating Unseen Parts in Dynamic 3D Human Reconstruction

In this paper, we present WonderHuman to reconstruct dynamic human avatars from a monocular video for high-fidelity novel view synthesis. Previous dynamic human avatar reconstruction methods typically require the input video to have full coverage of the observed human body. However, in daily practice, one typically has access to limited viewpoints, such as monocular front-view videos, making it a cumbersome task for previous methods to reconstruct the unseen parts of the human avatar. To tackle the issue, we present WonderHuman, which leverages 2D generative diffusion model priors to achieve high-quality, photorealistic reconstructions of dynamic human avatars from monocular videos, including accurate rendering of unseen body parts. Our approach introduces a Dual-Space Optimization technique, applying Score Distillation Sampling (SDS) in both canonical and observation spaces to ensure visual consistency and enhance realism in dynamic human reconstruction. Additionally, we present a View Selection strategy and Pose Feature Injection to enforce the consistency between SDS predictions and observed data, ensuring pose-dependent effects and higher fidelity in the reconstructed avatar. In the experiments, our method achieves SOTA performance in producing photorealistic renderings from the given monocular video, particularly for those challenging unseen parts. The project page and source code can be found at https://wyiguanw.github.io/WonderHuman/.

Zilong Wang, Zhiyang Dou, Yuan Liu, Cheng Lin, Xiao Dong, Yunhui Guo, Chenxu Zhang, Xin Li, Wenping Wang, Xiaohu Guo
arXiv:2502.01045 · cs.CV, cs.GR · submitted Feb 3, 2025 · updated Oct 5, 2025
abstract · pdf · html

add comment on HN
Also discussed: Feb 2025 (1 point, 0 comments)

This link (which is not underlined and so easy to miss, boooo) includes videos and more: https://wyiguanw.github.io/WonderHuman/
Not sure if there are similar techniques applied, but this definitely reminds of of the AI "microwave" filter that's been very popular on TikTok for the past few weeks. It creates equal numbers of impressive and hilarious 360 models from single images. One example: https://www.tiktok.com/@beccasthesia/video/74714552131157558...
Reminds me of that scene [1] from Enemy of the State where they "Rotate us 75 degrees around the vertical" in one of those "enhance, enhance" tropes.

[1] https://youtu.be/3EwZQddc3kY?t=7

Man, we all have been bluffed by this scene
Based on the video this looks great in the 200x200px previews typical for a paper but has substantial visual artifacts that become obvious at higher resolution.

Maybe there are straightforward ways to improve the results by spending more compute time. But as presented it's more "turn one picture into a convincing character in your Sims game", not "turn one picture into an animated character in your next TikTok video"

Looking forward to the next paper on this. Hopefully debuting the FaceBack app!