about
New 247k semantic similarity dataset: Image-Image, Image-Text, and Text-Text (arxiv.org)
2 points by danielcer on May 1, 2020 | hide | past | pdf | discuss on HN

In plain words: People rated how similar 267,095 pairs of images and captions are, adding the missing matches and mismatches the original image-caption set lacked. Training one model on both image-caption and caption-caption pairs showed the ratings can measure how learning across versus within each data type helps.

Abstract · Crisscrossed Captions: Extended Intramodal and Intermodal Semantic Similarity Judgments for MS-COCO

By supporting multi-modal retrieval training and evaluation, image captioning datasets have spurred remarkable progress on representation learning. Unfortunately, datasets have limited cross-modal associations: images are not paired with other images, captions are only paired with other captions of the same image, there are no negative associations and there are missing positive cross-modal associations. This undermines research into how inter-modality learning impacts intra-modality tasks. We address this gap with Crisscrossed Captions (CxC), an extension of the MS-COCO dataset with human semantic similarity judgments for 267,095 intra- and inter-modality pairs. We report baseline results on CxC for strong existing unimodal and multimodal models. We also evaluate a multitask dual encoder trained on both image-caption and caption-caption pairs that crucially demonstrates CxC's value for measuring the influence of intra- and inter-modality learning.

Zarana Parekh, Jason Baldridge, Daniel Cer, Austin Waters, Yinfei Yang
arXiv:2004.15020 · cs.CL · submitted Apr 30, 2020 · updated Mar 24, 2021
abstract · pdf · html · To be presented at EACL2021

add comment on HN