In plain words: A system finds people in surveillance video from a plain-language description of their height, clothing color, and gender. It first cuts each person out of the frame so background clutter is ignored, then matches the description, and it stayed accurate even in tough footage.
Abstract
A person is commonly described by attributes like height, build, cloth color, cloth type, and gender. Such attributes are known as soft biometrics. They bridge the semantic gap between human description and person retrieval in surveillance video. The paper proposes a deep learning-based linear filtering approach for person retrieval using height, cloth color, and gender. The proposed approach uses Mask R-CNN for pixel-wise person segmentation. It removes background clutter and provides precise boundary around the person. Color and gender models are fine-tuned using AlexNet and the algorithm is tested on SoftBioSearch dataset. It achieves good accuracy for person retrieval using the semantic query in challenging conditions.
Hiren Galiyawala, Kenil Shah, Vandit Gajjar, Mehul S. Raval
arXiv:1810.05080 · cs.CV · submitted Sep 24, 2018
abstract · pdf · 6 Pages, 6 Figures, Accepted to Semantic Person Retrieval in Surveillance Using Soft Biometrics challenge in Conjunction with AVSS-2018