In plain words: It makes fake image features for unseen combos like "red elephant" by reusing parts of seen concepts, adding task-specific variation at every layer. A classifier trained on them beats earlier approaches, more than doubling their accuracy when new and familiar combos are mixed.
Abstract · Task-Aware Feature Generation for Zero-Shot Compositional Learning
Visual concepts (e.g., red apple, big elephant) are often semantically compositional and each element of the compositions can be reused to construct novel concepts (e.g., red elephant). Compositional feature synthesis, which generates image feature distributions exploiting the semantic compositionality, is a promising approach to sample-efficient model generalization. In this work, we propose a task-aware feature generation (TFG) framework for compositional learning, which generates features of novel visual concepts by transferring knowledge from previously seen concepts. These synthetic features are then used to train a classifier to recognize novel concepts in a zero-shot manner. Our novel TFG design injects task-conditioned noise layer-by-layer, producing task-relevant variation at each level. We find the proposed generator design improves classification accuracy and sample efficiency. Our model establishes a new state of the art on three zero-shot compositional learning (ZSCL) benchmarks, outperforming the previous discriminative models by a large margin. Our model improves the performance of the prior arts by over 2x in the generalized ZSCL setting.
Xin Wang, Fisher Yu, Trevor Darrell, Joseph E. Gonzalez
arXiv:1906.04854 · cs.CV, cs.LG · submitted Jun 11, 2019 · updated Mar 23, 2020
abstract · pdf · html · 17 pages, 9 figures; substantial content updates with additional experiments