In plain words: It tests whether a human-readable concept truly affects a model's predictions by checking if they stay linked once other factors are fixed. A bet that grows as evidence builds up ranks concepts by importance while keeping false alarms controlled, which ordinary scores cannot promise.
Abstract
Recent works have extended notions of feature importance to semantic concepts that are inherently interpretable to the users interacting with a black-box predictive model. Yet, precise statistical guarantees, such as false positive rate and false discovery rate control, are needed to communicate findings transparently and to avoid unintended consequences in real-world scenarios. In this paper, we formalize the global (i.e., over a population) and local (i.e., for a sample) statistical importance of semantic concepts for the predictions of opaque models by means of conditional independence, which allows for rigorous testing. We use recent ideas of sequential kernelized independence testing (SKIT) to induce a rank of importance across concepts, and showcase the effectiveness and flexibility of our framework on synthetic datasets as well as on image classification tasks using several and diverse vision-language models.
Jacopo Teneggi, Jeremias Sulam
arXiv:2405.19146 · stat.ML, cs.LG · submitted May 29, 2024 · updated Oct 7, 2024
abstract · pdf · html