In plain words: Control theory's stability math, normally used to keep machines steady, is turned into a defense that provably blocks crafted input tricks against neural networks. Unlike usual defenses checked only by experiment, it gives a guaranteed bound, though only against a weaker attacker.
Abstract
Significant work is being done to develop the math and tools necessary to build provable defenses, or at least bounds, against adversarial attacks of neural networks. In this work, we argue that tools from control theory could be leveraged to aid in defending against such attacks. We do this by example, building a provable defense against a weaker adversary. This is done so we can focus on the mechanisms of control theory, and illuminate its intrinsic value.
Arash Rahnama, Andre T. Nguyen, Edward Raff
arXiv:1907.07732 · cs.CR, cs.LG · submitted Jul 17, 2019
abstract · pdf · html · 8 pages, 3 figures, AdvML'19: Workshop on Adversarial Learning Methods for Machine Learning and Data Mining at KDD