about
ADOP: Approximate Differentiable One-Pixel Point Rendering (arxiv.org)
45 points by rohithkp on Oct 16, 2021 | hide | past | pdf | 6 comments on HN

In plain words: It draws a scene's point cloud as one-pixel dots, then a neural network fills the gaps and shades the image, while also learning each photo's exposure and color balance. This lets it fix messy inputs and render over 100 million points in real time.

Abstract

In this paper we present ADOP, a novel point-based, differentiable neural rendering pipeline. Like other neural renderers, our system takes as input calibrated camera images and a proxy geometry of the scene, in our case a point cloud. To generate a novel view, the point cloud is rasterized with learned feature vectors as colors and a deep neural network fills the remaining holes and shades each output pixel. The rasterizer renders points as one-pixel splats, which makes it very fast and allows us to compute gradients with respect to all relevant input parameters efficiently. Furthermore, our pipeline contains a fully differentiable physically-based photometric camera model, including exposure, white balance, and a camera response function. Following the idea of inverse rendering, we use our renderer to refine its input in order to reduce inconsistencies and optimize the quality of its output. In particular, we can optimize structural parameters like the camera pose, lens distortions, point positions and features, and a neural environment map, but also photometric parameters like camera response function, vignetting, and per-image exposure and white balance. Because our pipeline includes photometric parameters, e.g.~exposure and camera response function, our system can smoothly handle input images with varying exposure and white balance, and generates high-dynamic range output. We show that due to the improved input, we can achieve high render quality, also for difficult input, e.g. with imperfect camera calibrations, inaccurate proxy geometry, or varying exposure. As a result, a simpler and thus faster deep neural network is sufficient for reconstruction. In combination with the fast point rasterization, ADOP achieves real-time rendering rates even for models with well over 100M points. https://github.com/darglein/ADOP

Darius Rückert, Linus Franke, Marc Stamminger
arXiv:2110.06635 · cs.CV, cs.GR · submitted Oct 13, 2021 · updated May 3, 2022
abstract · pdf · html

add comment on HN

Thank you, this gave me a very good introduction to the work!
So, if I understand correctly, they have created an algorithm that takes several colored point clouds (or photos + depth maps) of a scene and can generate renderings of the scene from novel viewpoints (or with different camera parameters, e.g. exposure or white balance). The results look truly impressive!

The whole rendering pipeline is differentiable, so they can backpropagate the ground truth error to all the unknowns in the pipeline. Those include values like camera pose parameters, point cloud texture and point positions, weights of a neural network that turns (potentially sparse) point cloud rasterizations into full HDR (high dynamic range) images, and tone mapping parameters like exposure.

(If I understand correctly, they start from something like a SLAM-based point cloud, but I haven't found that explained in glancing over the paper.)

Nitpick for the paper: A brief note on what HDR and LDR mean would be useful. The acronyms are not even expanded anywhere.

Structure from Motion (SfM) generated point clouds. Typical SLAM point clouds are much sparser so they can be run in real-time.
High and Low Dynamic Range
As I'm interested in the subject, I immediately think about the possibility of point cloud tech in a game development context. As far as I know, the big hurdle has always been dynamism like animation/deformation. Does this approach have any influence on that aspect?