In plain words: Instead of artists building sprite libraries by hand, it learns a set of reusable transparent patches and a network that places them on the canvas, with no labels. The result is a sparse, consistent set of parts that can be edited or analyzed directly.
Abstract
Artists and video game designers often construct 2D animations using libraries of sprites -- textured patches of objects and characters. We propose a deep learning approach that decomposes sprite-based video animations into a disentangled representation of recurring graphic elements in a self-supervised manner. By jointly learning a dictionary of possibly transparent patches and training a network that places them onto a canvas, we deconstruct sprite-based content into a sparse, consistent, and explicit representation that can be easily used in downstream tasks, like editing or analysis. Our framework offers a promising approach for discovering recurring visual patterns in image collections without supervision.
Dmitriy Smirnov, Michael Gharbi, Matthew Fisher, Vitor Guizilini, Alexei A. Efros, Justin Solomon
arXiv:2104.14553 · cs.CV · submitted Apr 29, 2021 · updated Oct 20, 2021
abstract · pdf · html · Accepted to NeurIPS 2021
It's possible to be generally intelligent without having any visual sense, so this approach may not help to develop General AI, but it could still be useful for many machine learning tasks.