about
Spatio-Temporal Action Graph Networks (arxiv.org)
3 points by sel1 on Oct 1, 2019 | hide | past | pdf | discuss on HN

In plain words: It builds a graph of how objects in a scene interact, learning separate patterns for their spatial arrangement and how it changes over time. Recognizing activities this way outperformed both plain scene-based models and earlier graph models.

Abstract

Events defined by the interaction of objects in a scene are often of critical importance; yet important events may have insufficient labeled examples to train a conventional deep model to generalize to future object appearance. Activity recognition models that represent object interactions explicitly have the potential to learn in a more efficient manner than those that represent scenes with global descriptors. We propose a novel inter-object graph representation for activity recognition based on a disentangled graph embedding with direct observation of edge appearance. We employ a novel factored embedding of the graph structure, disentangling a representation hierarchy formed over spatial dimensions from that found over temporal variation. We demonstrate the effectiveness of our model on the Charades activity recognition benchmark, as well as a new dataset of driving activities focusing on multi-object interactions with near-collision events. Our model offers significantly improved performance compared to baseline approaches without object-graph representations, or with previous graph-based models.

Roei Herzig, Elad Levi, Huijuan Xu, Hang Gao, Eli Brosh, Xiaolong Wang, Amir Globerson, Trevor Darrell
arXiv:1812.01233 · cs.CV · submitted Dec 4, 2018 · updated Sep 29, 2019
abstract · pdf · html · IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), 2019

add comment on HN