about
Progressive Relation Learning for Group Activity Recognition (arxiv.org)
2 points by sel1 on Aug 12, 2019 | hide | past | pdf | discuss on HN

In plain words: It builds a graph of how people in a video relate, then two agents trim it: one keeps the most telling frames, the other drops links that don't matter. They beat the usual approach of treating every person and frame equally on two benchmarks.

Abstract

Group activities usually involve spatiotemporal dynamics among many interactive individuals, while only a few participants at several key frames essentially define the activity. Therefore, effectively modeling the group-relevant and suppressing the irrelevant actions (and interactions) are vital for group activity recognition. In this paper, we propose a novel method based on deep reinforcement learning to progressively refine the low-level features and high-level relations of group activities. Firstly, we construct a semantic relation graph (SRG) to explicitly model the relations among persons. Then, two agents adopting policy according to two Markov decision processes are applied to progressively refine the SRG. Specifically, one feature-distilling (FD) agent in the discrete action space refines the low-level spatio-temporal features by distilling the most informative frames. Another relation-gating (RG) agent in continuous action space adjusts the high-level semantic graph to pay more attention to group-relevant relations. The SRG, FD agent, and RG agent are optimized alternately to mutually boost the performance of each other. Extensive experiments on two widely used benchmarks demonstrate the effectiveness and superiority of the proposed approach.

Guyue Hu, Bo Cui, Yuan He, Shan Yu
arXiv:1908.02948 · cs.CV · submitted Aug 8, 2019 · updated Mar 3, 2020
abstract · pdf · html · 8 pages; Accepted to CVPR2020; Supplementary Materials will appear on site <CVF Open Access>;

add comment on HN