about
Forecasting Human Dynamics from Static Images (2017) (arxiv.org)
1 point by seycombi on Apr 13, 2017 | hide | past | pdf | discuss on HN

In plain words: A system reads a single photo, guesses the person's pose in 2D, lifts it into 3D, then rolls out the next few poses. Trained on photos, videos, and motion-capture suits, it performs about as well as tools made for those tasks alone.

Abstract · Forecasting Human Dynamics from Static Images

This paper presents the first study on forecasting human dynamics from static images. The problem is to input a single RGB image and generate a sequence of upcoming human body poses in 3D. To address the problem, we propose the 3D Pose Forecasting Network (3D-PFNet). Our 3D-PFNet integrates recent advances on single-image human pose estimation and sequence prediction, and converts the 2D predictions into 3D space. We train our 3D-PFNet using a three-step training strategy to leverage a diverse source of training data, including image and video based human pose datasets and 3D motion capture (MoCap) data. We demonstrate competitive performance of our 3D-PFNet on 2D pose forecasting and 3D pose recovery through quantitative and qualitative results.

Yu-Wei Chao, Jimei Yang, Brian Price, Scott Cohen, Jia Deng
arXiv:1704.03432 · cs.CV · submitted Apr 11, 2017
abstract · pdf · html · Accepted in CVPR 2017

add comment on HN