In plain words: A system turns a photo into a 3D LEGO model by learning a compact code for shapes, predicting it from the image, then converting it into bricks. It is the first to make LEGO models from photos, and they were assembled with real bricks.
Abstract · Image2Lego: Customized LEGO Set Generation from Images
Although LEGO sets have entertained generations of children and adults, the challenge of designing customized builds matching the complexity of real-world or imagined scenes remains too great for the average enthusiast. In order to make this feat possible, we implement a system that generates a LEGO brick model from 2D images. We design a novel solution to this problem that uses an octree-structured autoencoder trained on 3D voxelized models to obtain a feasible latent representation for model reconstruction, and a separate network trained to predict this latent representation from 2D images. LEGO models are obtained by algorithmic conversion of the 3D voxelized model to bricks. We demonstrate first-of-its-kind conversion of photographs to 3D LEGO models. An octree architecture enables the flexibility to produce multiple resolutions to best fit a user's creative vision or design needs. In order to demonstrate the broad applicability of our system, we generate step-by-step building instructions and animations for LEGO models of objects and human faces. Finally, we test these automatically generated LEGO sets by constructing physical builds using real LEGO bricks.
Kyle Lennon, Katharina Fransen, Alexander O'Brien, Yumeng Cao, Matthew Beveridge, Yamin Arefeen, Nikhil Singh, Iddo Drori
arXiv:2108.08477 · cs.CV, cs.LG · submitted Aug 19, 2021
abstract · pdf · html · 9 pages, 10 figures
That reads to me like: "I'm not playing honey, this is work!" I sure hope the team doing this had a lot of fun :)
The algorithm itself is OK but doesn't appear all that novel to me. It looks like they're using ShapeNet (or similar dataset) to train a 3D voxel autoencoder which can predict geometry from an image. And then they convert those voxels to LEGO instructions.
For the first part, I've seen equally good papers before, for example by the ShapeNet team. For the second part, there's working explicit algorithms. So their main contribution appears to be to combine two working things into a (somewhat) useful whole.