In plain words: A trial-and-error learner picks which connections to add or remove in a graph, changing only the data's structure to trick a classifier while seeing just its predicted labels. These edits fooled models that learn from graph data on whole-graph and single-node classification.
Abstract
Deep learning on graph structures has shown exciting results in various applications. However, few attentions have been paid to the robustness of such models, in contrast to numerous research work for image or text adversarial attack and defense. In this paper, we focus on the adversarial attacks that fool the model by modifying the combinatorial structure of data. We first propose a reinforcement learning based attack method that learns the generalizable attack policy, while only requiring prediction labels from the target classifier. Also, variants of genetic algorithms and gradient methods are presented in the scenario where prediction confidence or gradients are available. We use both synthetic and real-world data to show that, a family of Graph Neural Network models are vulnerable to these attacks, in both graph-level and node-level classification tasks. We also show such attacks can be used to diagnose the learned classifiers.
Hanjun Dai, Hui Li, Tian Tian, Xin Huang, Lin Wang, Jun Zhu, Le Song
arXiv:1806.02371 · cs.LG, cs.CR, cs.SI, stat.ML · submitted Jun 6, 2018
abstract · pdf · html · to appear in ICML 2018