about
Primitive-Based Generation of Controllable and Editable 3D Semantic Scenes (arxiv.org)
3 points by PaulHoule on Jul 6, 2025 | hide | past | pdf | discuss on HN

In plain words: Instead of dense blocks, this tool builds 3D scenes from movable parts—objects as shapes and the ground as a flat grid—so pieces can be swapped or filled in. It made better scenes than block-based tools while using less memory and running faster.

Abstract · PrITTI: Primitive-based Generation of Controllable and Editable 3D Semantic Urban Scenes

Existing approaches to 3D semantic urban scene generation predominantly rely on voxel-based representations, which are bound by fixed resolution, challenging to edit, and memory-intensive in their dense form. In contrast, we advocate for a primitive-based paradigm where urban scenes are represented using compact, semantically meaningful 3D elements that are easy to manipulate and compose. To this end, we introduce PrITTI, a latent diffusion model that leverages vectorized object primitives and rasterized ground surfaces for generating diverse, controllable, and editable 3D semantic urban scenes. This hybrid representation yields a structured latent space that facilitates object- and ground-level manipulation. Experiments on KITTI-360 show that primitive-based representations unlock the full capabilities of diffusion transformers, achieving state-of-the-art 3D scene generation quality with lower memory requirements, faster inference, and greater editability than voxel-based methods. Beyond generation, PrITTI supports a range of downstream applications, including scene editing, inpainting, outpainting, and photo-realistic street-view synthesis. The source code and more results can be found at https://raniatze.github.io/pritti/.

Christina Ourania Tze, Daniel Dauner, Yiyi Liao, Dzmitry Tsishkou, Andreas Geiger
arXiv:2506.19117 · cs.CV · submitted Jun 23, 2025 · updated May 24, 2026
abstract · pdf · html · Accepted to CVPR 2026

add comment on HN