In plain words: A neural network learns to name people in everyday photos by picking up body cues like clothing and pose, not just faces, since faces are often hidden or blurry. It beat the previous best results on a large social-media photo collection.
Abstract
Recognising persons in everyday photos presents major challenges (occluded faces, different clothing, locations, etc.) for machine vision. We propose a convnet based person recognition system on which we provide an in-depth analysis of informativeness of different body cues, impact of training data, and the common failure modes of the system. In addition, we discuss the limitations of existing benchmarks and propose more challenging ones. Our method is simple and is built on open source and open data, yet it improves the state of the art results on a large dataset of social media photos (PIPA).
Seong Joon Oh, Rodrigo Benenson, Mario Fritz, Bernt Schiele
arXiv:1509.03502 · cs.CV · submitted Sep 11, 2015 · updated Sep 25, 2015
abstract · pdf · html · Accepted to ICCV 2015, revised