about
Wslln: Weakly Supervised Natural Language Localization Networks (arxiv.org)
2 points by sel1 on Sep 4, 2019 | hide | past | pdf | discuss on HN

In plain words: It finds where in a video an event described by a sentence happens, learning from just video-sentence pairs instead of the start and end times usual training needs. It scored best on two video collections, matching or beating methods given those time labels.

Abstract · WSLLN: Weakly Supervised Natural Language Localization Networks

We propose weakly supervised language localization networks (WSLLN) to detect events in long, untrimmed videos given language queries. To learn the correspondence between visual segments and texts, most previous methods require temporal coordinates (start and end times) of events for training, which leads to high costs of annotation. WSLLN relieves the annotation burden by training with only video-sentence pairs without accessing to temporal locations of events. With a simple end-to-end structure, WSLLN measures segment-text consistency and conducts segment selection (conditioned on the text) simultaneously. Results from both are merged and optimized as a video-sentence matching problem. Experiments on ActivityNet Captions and DiDeMo demonstrate that WSLLN achieves state-of-the-art performance.

Mingfei Gao, Larry S. Davis, Richard Socher, Caiming Xiong
arXiv:1909.00239 · cs.CV · submitted Aug 31, 2019
abstract · pdf · html · accepted by EMNLP2019

add comment on HN