about
Equivariant Neural Rendering (arxiv.org)
2 points by ramzyo on Nov 28, 2020 | hide | past | pdf | discuss on HN

In plain words: The model learns 3D scenes from plain photos by requiring its inner picture to shift and rotate exactly as a real scene would, without 3D labels. It renders new views in real time at about the same quality as models that take minutes.

Abstract

We propose a framework for learning neural scene representations directly from images, without 3D supervision. Our key insight is that 3D structure can be imposed by ensuring that the learned representation transforms like a real 3D scene. Specifically, we introduce a loss which enforces equivariance of the scene representation with respect to 3D transformations. Our formulation allows us to infer and render scenes in real time while achieving comparable results to models requiring minutes for inference. In addition, we introduce two challenging new datasets for scene representation and neural rendering, including scenes with complex lighting and backgrounds. Through experiments, we show that our model achieves compelling results on these datasets as well as on standard ShapeNet benchmarks.

Emilien Dupont, Miguel Angel Bautista, Alex Colburn, Aditya Sankar, Carlos Guestrin, Josh Susskind, Qi Shan
arXiv:2006.07630 · cs.CV, stat.ML · submitted Jun 13, 2020 · updated Dec 21, 2020
abstract · pdf · html · Add link to code

add comment on HN