about
Image2Mesh: Single Image 3D Reconstruction (arxiv.org)
97 points by cpheinrich on Feb 24, 2018 | hide | past | pdf | 15 comments on HN

In plain words: Unlike voxel blocks or point clouds, this system turns a photo into a mesh — a surface of triangles — by predicting a few numbers that bend and blend shapes. It rebuilds detailed objects from real and synthetic images without outlines or marked points older methods need.

Abstract · Image2Mesh: A Learning Framework for Single Image 3D Reconstruction

One challenge that remains open in 3D deep learning is how to efficiently represent 3D data to feed deep networks. Recent works have relied on volumetric or point cloud representations, but such approaches suffer from a number of issues such as computational complexity, unordered data, and lack of finer geometry. This paper demonstrates that a mesh representation (i.e. vertices and faces to form polygonal surfaces) is able to capture fine-grained geometry for 3D reconstruction tasks. A mesh however is also unstructured data similar to point clouds. We address this problem by proposing a learning framework to infer the parameters of a compact mesh representation rather than learning from the mesh itself. This compact representation encodes a mesh using free-form deformation and a sparse linear combination of models allowing us to reconstruct 3D meshes from single images. In contrast to prior work, we do not rely on silhouettes and landmarks to perform 3D reconstruction. We evaluate our method on synthetic and real-world datasets with very promising results. Our framework efficiently reconstructs 3D objects in a low-dimensional way while preserving its important geometrical aspects.

Jhony K. Pontes, Chen Kong, Sridha Sridharan, Simon Lucey, Anders Eriksson, Clinton Fookes
arXiv:1711.10669 · cs.CV · submitted Nov 29, 2017
abstract · pdf · html · 9 pages, 4 figures

add comment on HN

TLDR They use a neural net to process an image to find a similar known 3D model to the object in the image as well as parameters for how to deform that model to be like the object in the image, and also then perform a linear combination of some related meshes (this is a follow up to their prior work "Compact Model Representation for 3D Reconstruction"). Pretty exciting work, but still far from being able to robustly do Image2Mesh. I happen to have worked on something quite similar, if any of you are curious: https://deformnet-site.github.io/DeformNet-website/
Interesting, thanks for the link. I tried a few published methods for doing this a year or so and I found they were quite slow. As in, you'd rather just use a stereo camera to quickly generate a decent pointcloud than use a single view plus a network like these. What did you find your performance (ms/frame) was like?
With separate dedicated identifiers and matching reconstruction pipelines, while only caring about a limited number of things, as in our specializing in faces, we can reconstruct 25 to 30 independent people, their heads and shoulder tops only, in real time on a 3.4 GHz i7. Note the reconstruction is accurate only as far as it matters for FR, as in we tend to ignore hair. So the models are bald, but they have the frame to frame facial expressions for everyone. Here's a gif of a degraded lighting test at 70 fps with one head doing a 3D reconstruction independently performed each frame: https://www.dropbox.com/s/0jtee4ofh2vdeg0/BlakeHeadMeshes.gi... I'm just visualizing the face, but the entire head and shoulder tops are reconstructed.
Ah this is cool too! I assume no neural networks are involved here?

EDIT: why the changing colors in that gif?

The changing colors and grain are part of a degraded video while tracking the face test. The face detection and portions of the reconstruction are NNs.
Surprised this is getting any notice at all. I work at a facial recognition firm pioneering this technique 20 years ago. (www.CyberExtruder.com) The TLDR description by andreyk pretty much sums the process, as it originated long ago. It is very similar now, just with a team of very smart people iterating over the pipeline for so long has made it sophisticated to the point the TLDR description could be argued as too light.

I even made a 3D avatar API & service, now closed, from the technology 10 years ago. The twitter site is still around, if you want to see the quality of reconstruction possible: https://twitter.com/3DAvatarStore/media

No source code :(
No code yet, but promised: http://www.cs.cmu.edu/~chenk/
analogous to neural net super-resolution -- producing plausible hallucinations (in this case, 3D objects instead of more pixels)