In plain words: A system turns drawings into graphics code: a vision network guesses the shapes, then a second step turns that list into code with loops and variables. This lets it fix the network's mistakes, compare drawings by structure, and extend them past what was drawn.
Abstract
We introduce a model that learns to convert simple hand drawings into graphics programs written in a subset of \LaTeX. The model combines techniques from deep learning and program synthesis. We learn a convolutional neural network that proposes plausible drawing primitives that explain an image. These drawing primitives are like a trace of the set of primitive commands issued by a graphics program. We learn a model that uses program synthesis techniques to recover a graphics program from that trace. These programs have constructs like variable bindings, iterative loops, or simple kinds of conditionals. With a graphics program in hand, we can correct errors made by the deep network, measure similarity between drawings by use of similar high-level geometric structures, and extrapolate drawings. Taken together these results are a step towards agents that induce useful, human-readable programs from perceptual input.
Kevin Ellis, Daniel Ritchie, Armando Solar-Lezama, Joshua B. Tenenbaum
arXiv:1707.09627 · cs.AI · submitted Jul 30, 2017 · updated Oct 26, 2018
abstract · pdf · html