about
Learning to Infer Generative Template Programs for Visual Concepts (arxiv.org)
2 points by PaulHoule on Apr 6, 2024 | hide | past | pdf | discuss on HN

In plain words: The system learns to write recipe-like programs describing a visual concept's structure from a few examples, then uses them to make images or split objects into parts. Across 2D layouts, characters, and 3D shapes, it beat task-specific rivals and stayed competitive with single-domain specialists.

Abstract

People grasp flexible visual concepts from a few examples. We explore a neurosymbolic system that learns how to infer programs that capture visual concepts in a domain-general fashion. We introduce Template Programs: programmatic expressions from a domain-specific language that specify structural and parametric patterns common to an input concept. Our framework supports multiple concept-related tasks, including few-shot generation and co-segmentation through parsing. We develop a learning paradigm that allows us to train networks that infer Template Programs directly from visual datasets that contain concept groupings. We run experiments across multiple visual domains: 2D layouts, Omniglot characters, and 3D shapes. We find that our method outperforms task-specific alternatives, and performs competitively against domain-specific approaches for the limited domains where they exist.

R. Kenny Jones, Siddhartha Chaudhuri, Daniel Ritchie
arXiv:2403.15476 · cs.CV, cs.AI, cs.GR, cs.LG · submitted Mar 20, 2024 · updated Jun 9, 2024
abstract · pdf · html · ICML 2024; Project page: https://rkjones4.github.io/template.html

add comment on HN