about
Multi-Stream Single Shot Spatial-Temporal Action Detection (arxiv.org)
2 points by sel1 on Aug 26, 2019 | hide | past | pdf | discuss on HN

In plain words: A system that boxes actions in video by combining the current frame's picture and motion with short clips of past frames, all in one pass instead of the usual two-step search-then-check. It scored 71.3% on a standard action test, the best among one-pass detectors.

Abstract

We present a 3D Convolutional Neural Networks (CNNs) based single shot detector for spatial-temporal action detection tasks. Our model includes: (1) two short-term appearance and motion streams, with single RGB and optical flow image input separately, in order to capture the spatial and temporal information for the current frame; (2) two long-term 3D ConvNet based stream, working on sequences of continuous RGB and optical flow images to capture the context from past frames. Our model achieves strong performance for action detection in video and can be easily integrated into any current two-stream action detection methods. We report a frame-mAP of 71.30% on the challenging UCF101-24 actions dataset, achieving the state-of-the-art result of the one-stage methods. To the best of our knowledge, our work is the first system that combined 3D CNN and SSD in action detection tasks.

Pengfei Zhang, Yu Cao, Benyuan Liu
arXiv:1908.08178 · cs.CV · submitted Aug 22, 2019
abstract · pdf · html

add comment on HN