about
The many Shapley values for model explanation (arxiv.org)
2 points by sel1 on Aug 24, 2019 | hide | past | pdf | 1 comment on HN

In plain words: Shapley values split a prediction's credit among features, but the many ways to set them up give very different answers—sometimes even crediting features the model never uses. Baseline Shapley fixes the setup so one answer is the only one that satisfies the good rules.

Abstract

The Shapley value has become a popular method to attribute the prediction of a machine-learning model on an input to its base features. The use of the Shapley value is justified by citing [16] showing that it is the \emph{unique} method that satisfies certain good properties (\emph{axioms}). There are, however, a multiplicity of ways in which the Shapley value is operationalized in the attribution problem. These differ in how they reference the model, the training data, and the explanation context. These give very different results, rendering the uniqueness result meaningless. Furthermore, we find that previously proposed approaches can produce counterintuitive attributions in theory and in practice---for instance, they can assign non-zero attributions to features that are not even referenced by the model. In this paper, we use the axiomatic approach to study the differences between some of the many operationalizations of the Shapley value for attribution, and propose a technique called Baseline Shapley (BShap) that is backed by a proper uniqueness result. We also contrast BShap with Integrated Gradients, another extension of Shapley value to the continuous setting.

Mukund Sundararajan, Amir Najmi
arXiv:1908.08474 · cs.AI, cs.LG, econ.TH · submitted Aug 22, 2019 · updated Feb 7, 2020
abstract · pdf · html · 9 pages

add comment on HN

> This is the crux of the matter: In taking the model seriously (even "literally"), BS evaluates it on points in feature space that may never occur in practice, and may not even be realizable.

Is it April 1?