In plain words: A system pastes a person into a photo and works out poses that fit the surroundings, like sitting on a chair, learning this by re-posing people in video clips. It made more realistic-looking people and more natural scene interactions than earlier tools.
Abstract
We study the problem of inferring scene affordances by presenting a method for realistically inserting people into scenes. Given a scene image with a marked region and an image of a person, we insert the person into the scene while respecting the scene affordances. Our model can infer the set of realistic poses given the scene context, re-pose the reference person, and harmonize the composition. We set up the task in a self-supervised fashion by learning to re-pose humans in video clips. We train a large-scale diffusion model on a dataset of 2.4M video clips that produces diverse plausible poses while respecting the scene context. Given the learned human-scene composition, our model can also hallucinate realistic people and scenes when prompted without conditioning and also enables interactive editing. A quantitative evaluation shows that our method synthesizes more realistic human appearance and more natural human-scene interactions than prior work.
Sumith Kulal, Tim Brooks, Alex Aiken, Jiajun Wu, Jimei Yang, Jingwan Lu, Alexei A. Efros, Krishna Kumar Singh
arXiv:2304.14406 · cs.CV, cs.AI, cs.GR, cs.LG · submitted Apr 27, 2023
abstract · pdf · html · CVPR 2023. Project page with code: https://sumith1896.github.io/affordance-insertion/