In plain words: They compute how a class score changes as each pixel changes, showing what a classifier learned: one builds an image that maximizes a class score, the other marks pixels that matter for a class. Those marks alone can segment objects without pixel labels.
Abstract · Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
This paper addresses the visualisation of image classification models, learnt using deep Convolutional Networks (ConvNets). We consider two visualisation techniques, based on computing the gradient of the class score with respect to the input image. The first one generates an image, which maximises the class score [Erhan et al., 2009], thus visualising the notion of the class, captured by a ConvNet. The second technique computes a class saliency map, specific to a given image and class. We show that such maps can be employed for weakly supervised object segmentation using classification ConvNets. Finally, we establish the connection between the gradient-based ConvNet visualisation methods and deconvolutional networks [Zeiler et al., 2013].
Karen Simonyan, Andrea Vedaldi, Andrew Zisserman
arXiv:1312.6034 · cs.CV · submitted Dec 20, 2013 · updated Apr 19, 2014
abstract · pdf · html