In plain words: A tool reads a network's on/off neuron patterns and turns them into simple conditions—on the input or inside the network—that guarantee a given output class. On digit-recognition and aircraft-avoidance networks, these conditions explained predictions, proved robustness, shortened proofs, and shrank the network.
Abstract · Property Inference for Deep Neural Networks
We present techniques for automatically inferring formal properties of feed-forward neural networks. We observe that a significant part (if not all) of the logic of feed forward networks is captured in the activation status ('on' or 'off') of its neurons. We propose to extract patterns based on neuron decisions as preconditions that imply certain desirable output property e.g., the prediction being a certain class. We present techniques to extract input properties, encoding convex predicates on the input space that imply given output properties and layer properties, representing network properties captured in the hidden layers that imply the desired output behavior. We apply our techniques on networks for the MNIST and ACASXU applications. Our experiments highlight the use of the inferred properties in a variety of tasks, such as explaining predictions, providing robustness guarantees, simplifying proofs, and network distillation.
Divya Gopinath, Hayes Converse, Corina S. Pasareanu, Ankur Taly
arXiv:1904.13215 · cs.LG, cs.AI, cs.FL · submitted Apr 29, 2019 · updated Sep 10, 2020
abstract · pdf · html · Errata: This version updates the ASE'19 conference version by correcting the definition of the three properties that were checked for ACASXU