about
The (Un)reliability of saliency methods [pdf] (arxiv.org)
1 point by stablemap on Nov 6, 2017 | hide | past | pdf | discuss on HN

In plain words: Saliency methods highlight which parts of an input drive a neural network's prediction. Adding a constant shift to the input — a change that leaves the prediction untouched — makes many of them point to the wrong parts, so they should mirror the model's own sensitivity.

Abstract · The (Un)reliability of saliency methods

Saliency methods aim to explain the predictions of deep neural networks. These methods lack reliability when the explanation is sensitive to factors that do not contribute to the model prediction. We use a simple and common pre-processing step ---adding a constant shift to the input data--- to show that a transformation with no effect on the model can cause numerous methods to incorrectly attribute. In order to guarantee reliability, we posit that methods should fulfill input invariance, the requirement that a saliency method mirror the sensitivity of the model with respect to transformations of the input. We show, through several examples, that saliency methods that do not satisfy input invariance result in misleading attribution.

Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T. Schütt, Sven Dähne, Dumitru Erhan, Been Kim
arXiv:1711.00867 · stat.ML, cs.LG · submitted Nov 2, 2017
abstract · pdf · html

add comment on HN