about
Image Captioning with Unseen Objects (arxiv.org)
1 point by sel1 on Aug 4, 2019 | hide | past | pdf | discuss on HN

In plain words: A captioning system that spots objects it never saw in training, then fills them into sentence templates. Unlike captioners that only name objects their recognizer was trained on, it captions images with new objects and showed promising results on a standard photo set.

Abstract

Image caption generation is a long standing and challenging problem at the intersection of computer vision and natural language processing. A number of recently proposed approaches utilize a fully supervised object recognition model within the captioning approach. Such models, however, tend to generate sentences which only consist of objects predicted by the recognition models, excluding instances of the classes without labelled training examples. In this paper, we propose a new challenging scenario that targets the image captioning problem in a fully zero-shot learning setting, where the goal is to be able to generate captions of test images containing objects that are not seen during training. The proposed approach jointly uses a novel zero-shot object detection model and a template-based sentence generator. Our experiments show promising results on the COCO dataset.

Berkan Demirel, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis
arXiv:1908.00047 · cs.CV · submitted Jul 31, 2019
abstract · pdf · html · To appear in British Machine Vision Conference (BMVC) 2019

add comment on HN