In plain words: One network reads long-term motion tracking to tell actions apart, while another looks at recent 3D video frames for fine detail; their judgments are combined to classify the movement. Together they beat the best previous systems by 14% on standard test videos.
Abstract
The recognition of actions from video sequences has many applications in health monitoring, assisted living, surveillance, and smart homes. Despite advances in sensing, in particular related to 3D video, the methodologies to process the data are still subject to research. We demonstrate superior results by a system which combines recurrent neural networks with convolutional neural networks in a voting approach. The gated-recurrent-unit-based neural networks are particularly well-suited to distinguish actions based on long-term information from optical tracking data; the 3D-CNNs focus more on detailed, recent information from video data. The resulting features are merged in an SVM which then classifies the movement. In this architecture, our method improves recognition rates of state-of-the-art methods by 14% on standard data sets.
Rui Zhao, Haider Ali, Patrick van der Smagt
arXiv:1703.09783 · cs.CV, cs.LG · submitted Mar 22, 2017 · updated Oct 2, 2018
abstract · pdf · html · Published in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)