about
Contextual RNN-GANs for Abstract Reasoning Diagram Generation (arxiv.org)
2 points by MichaelBurge on Jan 14, 2017 | hide | past | pdf | 1 comment on HN

In plain words: A picture generator and its checker both remember earlier frames so the generator can draw the next image in an abstract reasoning puzzle. It matched 10th-grade humans on these puzzles, though college students did better, and improved video frame prediction.

Abstract

Understanding, predicting, and generating object motions and transformations is a core problem in artificial intelligence. Modeling sequences of evolving images may provide better representations and models of motion and may ultimately be used for forecasting, simulation, or video generation. Diagrammatic Abstract Reasoning is an avenue in which diagrams evolve in complex patterns and one needs to infer the underlying pattern sequence and generate the next image in the sequence. For this, we develop a novel Contextual Generative Adversarial Network based on Recurrent Neural Networks (Context-RNN-GANs), where both the generator and the discriminator modules are based on contextual history (modeled as RNNs) and the adversarial discriminator guides the generator to produce realistic images for the particular time step in the image sequence. We evaluate the Context-RNN-GAN model (and its variants) on a novel dataset of Diagrammatic Abstract Reasoning, where it performs competitively with 10th-grade human performance but there is still scope for interesting improvements as compared to college-grade human performance. We also evaluate our model on a standard video next-frame prediction task, achieving improved performance over comparable state-of-the-art.

Arnab Ghosh, Viveka Kulharia, Amitabha Mukerjee, Vinay Namboodiri, Mohit Bansal
arXiv:1609.09444 · cs.CV, cs.AI, cs.LG · submitted Sep 29, 2016 · updated Dec 6, 2016
abstract · pdf · html · To Appear in AAAI-17 and NIPS Workshop on Adversarial Training

add comment on HN

A dogs classifier can be thought of as training a predicate DOG(X), where X ranges over image fragments. Is there any research on training n-ary or nested relations?

I can already train something like OVER(CUP(X), TABLE(X)) as an ordinary classifier. But OVER() wouldn't generalize to predicates other than CUP or TABLE, and it couldn't take independent image fragments: OVER(CUP(X), TABLE(Y)), where X < Z and Y < Z, where < means (rectangular?) subset.

I'd like to train a GAN or similar to generate an image that satisfies arbitrary formulas in predicate logic, with a dataset of images that are tagged with formulas. Anyone have any ideas for that?