about
Improving NeRF quality by progressive camera placement (arxiv.org)
33 points by PaulHoule on Sep 10, 2023 | hide | past | pdf | 13 comments on HN

In plain words: A program picks where to put each new camera so the photos teach a 3D scene model to render any viewpoint cleanly, even in cluttered rooms. It beats fixed camera setups and earlier camera-picking tricks at making free-viewpoint views look good.

Abstract · Improving NeRF Quality by Progressive Camera Placement for Unrestricted Navigation in Complex Environments

Neural Radiance Fields, or NeRFs, have drastically improved novel view synthesis and 3D reconstruction for rendering. NeRFs achieve impressive results on object-centric reconstructions, but the quality of novel view synthesis with free-viewpoint navigation in complex environments (rooms, houses, etc) is often problematic. While algorithmic improvements play an important role in the resulting quality of novel view synthesis, in this work, we show that because optimizing a NeRF is inherently a data-driven process, good quality data play a fundamental role in the final quality of the reconstruction. As a consequence, it is critical to choose the data samples -- in this case the cameras -- in a way that will eventually allow the optimization to converge to a solution that allows free-viewpoint navigation with good quality. Our main contribution is an algorithm that efficiently proposes new camera placements that improve visual quality with minimal assumptions. Our solution can be used with any NeRF model and outperforms baselines and similar work.

Georgios Kopanas, George Drettakis
arXiv:2309.00014 · cs.CV, cs.GR, eess.IV · submitted Aug 24, 2023 · updated Sep 4, 2023
abstract · pdf · html

add comment on HN

I'm still thinking that a 3D video recording would be the best way to gather this information.

Don't worry about taking stills or optimal placement of the camera for stills, because all the information you need is captured in the process of taking a decent 3D video scan of each room. And each frame of that 3D video captures important information not only from the objects it's closest to, but also of the objects across the room, including which objects may partially or completely occlude other objects.

The geometry of still images are superior to that of videos, assuming proper gear. Thus, optimal frames taken with a good camera, is better than some video capture.
I wonder if techniques like this could work for Gaussian Splattering too?

Also, they mention AR headsets and Drones in the conclusion, which would both be fantastic use cases. High quality inspections could be done by allowing the drone to capture the absolute best angles just-in-time rather than a predetermined path

Most techniques for nerfs also work for the newer variants like guassian splatting.

It's the same fundamental process - optimizing a 3D representation of a scene using gradient descent. Neural networks are a very general-purpose way to represent anything, but that comes at a cost. Using a more 3D-specific representation improves performance.

... that's exactly what I find so exciting about this field, a factor of 3 here and a factor of 5 here and pretty soon I can go out with my lightfield camera and make a 3D model of your place of business to put it into the "metaverse" which you could never afford to do the way they make video games.
The authors wrote the GS paper
Good stuff coming from them and that lab!

From a researcher's personal page[1], they're working on cashing in that GS stuff: estimating pose dynamics by regularizing Gaussians’ motion and rotation with local-rigidity constraints [2]. The videos are impressive[3], and seem to be bringing in the 4th dimension.

Twitter is here[4]

[1] https://grgkopanas.github.io/

[2] https://grgkopanas.github.io/2023/08/24/dynamic_3d_gaussian....

[3] https://dynamic3dgaussians.github.io/

[4] https://twitter.com/GKopanas

"nerf" here meaning (as near as I can tell)

> view synthesis and 3D reconstruction for rendering

Yep, “neural radiance fields”, effectively neural networks can be trained to model how light is emitted and moves in a volume of space to, for instance, take a few pictures of a scene and then be able the synthesize views of the scene. If it can be practically sped up it would be a way to make VR models of spaces without specialist talent in 3-d modeling.
Gaussian splatting is an interesting take on the technique that seems to speed up fitting and rendering, see this article from a couple days ago: https://news.ycombinator.com/item?id=37415478
Yep, there is all sorts of research on that problem.
My understanding is that NeRF¹ is just one (very interesting) approach to this. Another approach would be "Multi-View Stereo (MVS)", which you may be familiar with via software like COLMAP.

¹ https://en.wikipedia.org/wiki/Neural_radiance_field

They use NeRF to mean a radiance field that can be sampled and trained using a differentiable process.