about
IM2CAD – Reconstruct a scene that is as similar as possible to a photograph (arxiv.org)
86 points by EvgeniyZh on Aug 27, 2016 | hide | past | pdf | 6 comments on HN

In plain words: From one room photo, the system picks furniture 3D models from a database and nudges their position and size until a rendering of the whole scene matches the picture, so objects correctly block each other. It rebuilds rooms automatically and improves on scene-understanding tests.

Abstract · IM2CAD

Given a single photo of a room and a large database of furniture CAD models, our goal is to reconstruct a scene that is as similar as possible to the scene depicted in the photograph, and composed of objects drawn from the database. We present a completely automatic system to address this IM2CAD problem that produces high quality results on challenging imagery from interior home design and remodeling websites. Our approach iteratively optimizes the placement and scale of objects in the room to best match scene renderings to the input photo, using image comparison metrics trained via deep convolutional neural nets. By operating jointly on the full scene at once, we account for inter-object occlusions. We also show the applicability of our method in standard scene understanding benchmarks where we obtain significant improvement.

Hamid Izadinia, Qi Shan, Steven M. Seitz
arXiv:1608.05137 · cs.CV · submitted Aug 18, 2016 · updated Apr 24, 2017
abstract · pdf · html · To appear at CVPR 2017

add comment on HN

Besides being a technological improvement this paper is really well written and readable.

https://arxiv.org/pdf/1608.05137v1.pdf

Wow.

(I don't have any idea of what I'm talking about but) it occurs to me that if a robot does this for the frames of its video input with a regular camera, then on static environments the output of SLAM would be great.

Also, just predict what you will see after you move from the CAD scene, move, compare the actual new image with the predicted one, and dedicate most computing resources to what differs the most - now you have a robot with attention to unrecognized objects!

I wonder how it could be used to efficiently generate maps for Counter-Strike (and other similar games) ;) Like pictures of high school...
This! I rememebr back in 00' we were always talking with friends about creating such map (school, neighbourhood, etc.) because it seemed funny to play on such map. Unfortunately nobody was skilled enough to make such map...
I did this when I was young! Except I did it for Quake 1.

I had great fun re-constructing various parts of my high school. I learnt about the importance of scale, how to make levels "interesting" by making some desks fall over, and had great fun reconstructing the swimming pool, gym, and metalworking facilities.

When I showed people they were always concerned because there was a gun on the screen and I was walking around my high school. I lacked the ability to get rid of the gun, so I had to explain it every time. Some of the more stupid teachers couldn't get past the gun and see it for what it was (a student interested in programming - not a cry for help) but I ended up just not telling them anything and taking it to people with brains.

Later on I started re-constructing the Titanic from the original schematics, though I never got very far with this, because my machine only had 64MB of ram, and I used it up pretty quickly making complicated geometry.

Good fun.

The development hopefully goes further to replace moving objects: humans, cars etc. for rebuilding 3D scenes from video for VR.