about
Google Brain: Concept Activation Vectors (arxiv.org)
4 points by bra-ket on Apr 11, 2018 | hide | past | pdf | discuss on HN

In plain words: Instead of highlighting pixels, this method turns a human idea like "stripes" into a direction inside the network's internal state and measures how much a prediction shifts along it. That yields a score for how much a concept matters, shown on photos and medical images.

Abstract · Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)

The interpretation of deep learning models is a challenge due to their size, complexity, and often opaque internal state. In addition, many systems, such as image classifiers, operate on low-level features rather than high-level concepts. To address these challenges, we introduce Concept Activation Vectors (CAVs), which provide an interpretation of a neural net's internal state in terms of human-friendly concepts. The key idea is to view the high-dimensional internal state of a neural net as an aid, not an obstacle. We show how to use CAVs as part of a technique, Testing with CAVs (TCAV), that uses directional derivatives to quantify the degree to which a user-defined concept is important to a classification result--for example, how sensitive a prediction of "zebra" is to the presence of stripes. Using the domain of image classification as a testing ground, we describe how CAVs may be used to explore hypotheses and generate insights for a standard image classification network as well as a medical application.

Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, Rory Sayres
arXiv:1711.11279 · stat.ML · submitted Nov 30, 2017 · updated Jun 7, 2018
abstract · pdf · html

add comment on HN
Also discussed: Feb 2019 (1 point, 0 comments)