about
Training Differentiable Models by Constraining Their Explanations (arxiv.org)
36 points by makmanalp on Mar 13, 2017 | hide | past | pdf | 5 comments on HN

In plain words: It watches which input features a prediction leans on—by nudging each input and seeing how the output shifts—and penalizes the wrong ones, guided by experts or on its own. Models trained this way generalized much better than usual when test data differed from training.

Abstract · Right for the Right Reasons: Training Differentiable Models by Constraining their Explanations

Neural networks are among the most accurate supervised learning methods in use today, but their opacity makes them difficult to trust in critical applications, especially when conditions in training differ from those in test. Recent work on explanations for black-box models has produced tools (e.g. LIME) to show the implicit rules behind predictions, which can help us identify when models are right for the wrong reasons. However, these methods do not scale to explaining entire datasets and cannot correct the problems they reveal. We introduce a method for efficiently explaining and regularizing differentiable models by examining and selectively penalizing their input gradients, which provide a normal to the decision boundary. We apply these penalties both based on expert annotation and in an unsupervised fashion that encourages diverse models with qualitatively different decision boundaries for the same classification problem. On multiple datasets, we show our approach generates faithful explanations and models that generalize much better when conditions differ between training and test.

Andrew Slavin Ross, Michael C. Hughes, Finale Doshi-Velez
arXiv:1703.03717 · cs.LG, cs.AI, stat.ML · submitted Mar 10, 2017 · updated May 25, 2017
abstract · pdf · html

add comment on HN

someone want to do us all a favor and define what an explanation is in this context
That explains LIME, an older paper that's not the one being discussed here (but is referenced)
It also explains what an explanation is in this context (which is what was asked): a local linear approximation of the model. Additionally it has a diagram which is nice. Obviously it's not the one being discussed here though -- I'd hardly be adding useful information if I just linked to the submission again as a reply.
Yeah, so local linear approximations are what we and LIME are using as explanations, but it's not what an explanation is generally.

In the paper we do define an explanation as basically any artifact that "provides reliable information about the model’s implicit decision rules for a given prediction." It's kind of a rough and over-general definition, but it gets to the idea that explanations can be partial. All we want to do is turn a completely black-box model into something slightly more transparent.

Ideally, we could have explanations that were at a higher level of abstraction, e.g. "this image is a picture of a husky and not a wolf because of the shape of the nose and the color of the coat," but a neural network has no idea what "nose" and "coat" means. Sometimes its intermediate layers will end up corresponding to meaningful abstract concepts like that, but not always.