about
Large-scale representation learning from visually grounded untranscribed speech (arxiv.org)
2 points by sel1 on Sep 21, 2019 | hide | past | pdf | discuss on HN

In plain words: By automatically turning image captions into varied spoken audio, this system trains a network to place matching sounds and pictures close together in a shared space. It retrieves the right image among its top 10 guesses 49.5% of the time, up from 29.6%.

Abstract

Systems that can associate images with their spoken audio captions are an important step towards visually grounded language learning. We describe a scalable method to automatically generate diverse audio for image captioning datasets. This supports pretraining deep networks for encoding both audio and images, which we do via a dual encoder that learns to align latent representations from both modalities. We show that a masked margin softmax loss for such models is superior to the standard triplet loss. We fine-tune these models on the Flickr8k Audio Captions Corpus and obtain state-of-the-art results---improving recall in the top 10 from 29.6% to 49.5%. We also obtain human ratings on retrieval outputs to better assess the impact of incidentally matching image-caption pairs that were not associated in the data, finding that automatic evaluation substantially underestimates the quality of the retrieved results.

Gabriel Ilharco, Yuan Zhang, Jason Baldridge
arXiv:1909.08782 · cs.CV, cs.CL, cs.SD, eess.AS · submitted Sep 19, 2019
abstract · pdf · html

add comment on HN