about
Provably Robust Deep Learning (arxiv.org)
2 points by thomasahle on Jun 13, 2019 | hide | past | pdf | discuss on HN

In plain words: Randomized smoothing makes a classifier provably hard to fool by adding noise to its inputs; here it is also trained on tricked examples using an attack built for noisy classifiers. It beats every earlier provably robust classifier by a wide margin on two image benchmarks.

Abstract · Provably Robust Deep Learning via Adversarially Trained Smoothed Classifiers

Recent works have shown the effectiveness of randomized smoothing as a scalable technique for building neural network-based classifiers that are provably robust to $\ell_2$-norm adversarial perturbations. In this paper, we employ adversarial training to improve the performance of randomized smoothing. We design an adapted attack for smoothed classifiers, and we show how this attack can be used in an adversarial training setting to boost the provable robustness of smoothed classifiers. We demonstrate through extensive experimentation that our method consistently outperforms all existing provably $\ell_2$-robust classifiers by a significant margin on ImageNet and CIFAR-10, establishing the state-of-the-art for provable $\ell_2$-defenses. Moreover, we find that pre-training and semi-supervised learning boost adversarially trained smoothed classifiers even further. Our code and trained models are available at http://github.com/Hadisalman/smoothing-adversarial .

Hadi Salman, Greg Yang, Jerry Li, Pengchuan Zhang, Huan Zhang, Ilya Razenshteyn, Sebastien Bubeck
arXiv:1906.04584 · cs.LG, cs.CR, stat.ML · submitted Jun 9, 2019 · updated Jan 10, 2020
abstract · pdf · html · Spotlight at the 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada; 9 pages main text; 31 pages total

add comment on HN