about
Where Do Models Go Wrong? Parameter-Space Saliency Maps for Explainability (arxiv.org)
84 points by lnyan on Aug 4, 2021 | hide | past | pdf | 6 comments on HN

In plain words: Instead of highlighting input pixels, this method finds which of a network's internal settings cause a wrong answer. Samples that trip the same settings are semantically similar, and removing or retraining just those settings fixes other mistakes of the same kind.

Abstract · Where do Models go Wrong? Parameter-Space Saliency Maps for Explainability

Conventional saliency maps highlight input features to which neural network predictions are highly sensitive. We take a different approach to saliency, in which we identify and analyze the network parameters, rather than inputs, which are responsible for erroneous decisions. We find that samples which cause similar parameters to malfunction are semantically similar. We also show that pruning the most salient parameters for a wrongly classified sample often improves model behavior. Furthermore, fine-tuning a small number of the most salient parameters on a single sample results in error correction on other samples that are misclassified for similar reasons. Based on our parameter saliency method, we also introduce an input-space saliency technique that reveals how image features cause specific network components to malfunction. Further, we rigorously validate the meaningfulness of our saliency maps on both the dataset and case-study levels.

Roman Levin, Manli Shu, Eitan Borgnia, Furong Huang, Micah Goldblum, Tom Goldstein
arXiv:2108.01335 · cs.CV, cs.LG · submitted Aug 3, 2021 · updated Oct 10, 2022
abstract · pdf · html

add comment on HN

Kudos to authors for such detailed work. I am yet to go through in detail but the way they have approached the problem certainly is refreshing. Any thoughts on how to extend this to non-vision architectures?
Thanks!

The parameter saliency approach can be naturally generalized to any architecture, not just vision. The changes that need to be made: 1. Aggregation. Currently, our code aggregates saliency of the model parameters by averaging on the conv filter level. We used that because in the literature filters have been shown to be interpretable. However, no aggregation can be used and the saliency profile can be computed on individual parameter level allowing for any architecture. 2. Loss. The loss is also not limited to classification losses, any other loss function can be used, e.g. metric learning or regression.

These should be fairly simple modifications of the code. Happy to help if needed!

I have benn working so so much with explainability. I think the main thing to know when trying to extend this to non vision (which I assume to be standard feature engineering based ml, let's ignore texts for a second) is that all the explainability is in the features. You have to design features that make sense. It really doesn't matter if you have the best explainability method in the world if that perfect explainability method retrasces the models prediction to features nobody understands.
Well, I agree to your point that the explainability is only as good as the features used in the model. Yet it is important to attribute the correct set of features (along with quantification) for a given instance. The SHAP framework does that to a good extent but it's focus is not towards helping with identification of model parameter related issues. This work here seems to focus on the later and hence I found it a bit different and a refreshing perspective towards the explainability aspect.
Oh my goodness I misread this as “modals” and was expecting a complex academic treatise on dialog boxes.
Reading the first part of the titles, I thought maybe it is about poor career choices of fashion models...